Data synchronization verification method and apparatus, and electronic device
By receiving and comparing the characteristic information of different data sources during the data synchronization process, data consistency is quickly determined, and data loss phenomenon during data synchronization process and the low efficiency of existing synchronization verification methods is solved, and efficient and accurate data synchronization verification is achieved.
Patent Information
- Application Number
- PCT/CN2024/121974
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-15
- Filing Date
- 2024-09-27
- Publication Date
- 2025-06-19
AI Technical Summary
Data loss may occur during data synchronization, resulting in inaccurate data analysis after synchronization. The existing technology synchronization verification methods take a long time, occupy a lot of resources, have large delays, low verification efficiency, and are high in artificial verification costs and large delays.
A data synchronization verification method is provided, by receiving characteristic information of the first and second data of different data sources, the consistency of the current batch of data is determined, and if it is consistent, the verification is determined to pass. The method includes a receiving module, a consistency determination module and a calibration result determination module, which can perform data synchronization verification quickly and accurately.
Accurate verification of data synchronization quality with fewer resources is achieved, reducing labor costs, improving verification efficiency and accuracy, and reducing verification time and delay.
Smart Images

Figure CN2024121974_19062025_PF_FP_ABST
Abstract
Description
Data synchronization verification method and device, and electronic equipment
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 15, 2023, with application number 202311733894.0 and application name “Data synchronization verification method and device, electronic device, storage medium”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to, but is not limited to, the field of data processing technology, and specifically to a data synchronization verification method and device, and an electronic device. Background Art
[0003] With the rapid development of the internet, massive amounts of data from heterogeneous data sources are generated daily, resulting in tens of thousands of data synchronization tasks. If data loss occurs during the synchronization process, it can impact post-synchronization data analysis and hinder the re-launch of production data. Accurately verifying the quality of data synchronization with minimal resources has become a pressing issue. Technical issues
[0004] Data loss may occur during the data synchronization process, which will have an adverse effect on the data analysis after synchronization. In the related art, some synchronization tasks lack the task of verifying the quality of data synchronization, and the other part of the synchronization tasks will configure a synchronization verification task to determine whether data loss occurs or perform manual verification. For example, the synchronization verification task can be a task that performs synchronization verification at the level of a single data item, and performing such synchronization verification on a large amount of heterogeneous data sources has problems such as being time-consuming, occupying too many resources, having large delays, and having low verification efficiency. Manual verification will also cause a lot of manpower costs and large delays, and data synchronization will be performed every day, which will bring a lot of repetitive work. Whether it is the synchronization verification process after missing data synchronization or the delayed discovery of data loss, it will result in the use of incomplete data for data analysis in an uncertain time period, affecting the accuracy of the data analysis results. Technical Solutions
[0005] The following is a summary of the subject matter described in detail herein. This summary is not intended to limit the scope of the claims.
[0006] In a first aspect, the present application provides a data synchronization verification method, the method comprising:
[0007] receiving first characteristic information of first data and second characteristic information of second data, wherein the first data and the second data come from different data sources during data synchronization and include different batches of data;
[0008] determining whether the first data of the current batch is consistent with the second data according to the first characteristic information of the first data of the current batch and the second characteristic information of the second data of the current batch;
[0009] When the first data and the second data of the current batch are consistent, it is determined that the result of the data synchronization verification of the current batch is verification passed.
[0010] In a second aspect, the present application provides a data synchronization verification device, which includes:
[0011] a receiving module configured to receive first characteristic information of first data and second characteristic information of second data, wherein the first data and the second data are respectively from different data sources in a data synchronization process and respectively include different batches of data;
[0012] a consistency determination module configured to determine whether the first data of the current batch is consistent with the second data based on the first characteristic information of the first data of the current batch and the second characteristic information of the second data of the current batch;
[0013] The verification result determination module is configured to determine that the data synchronization verification result of the current batch is verification passed when the first data and the second data of the current batch are consistent.
[0014] In a third aspect, the present application provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores one or more computer programs executable by the at least one processor, and one or more of the computer programs are executed by the at least one processor so that the at least one processor can execute the above-mentioned data synchronization verification method.
[0015] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the above-mentioned data synchronization verification method when executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings are used to provide a further understanding of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. The above and other features and advantages will become more apparent to those skilled in the art by describing the detailed exemplary embodiments with reference to the accompanying drawings. In the accompanying drawings:
[0017] FIG1 is a flow chart of a data synchronization verification method provided by an embodiment of the present application;
[0018] FIG2 is a schematic diagram of an application scenario of a data synchronization verification method provided in an embodiment of the present application;
[0019] FIG3 is a schematic diagram of an application scenario of a data synchronization verification method provided in an embodiment of the present application;
[0020] FIG4 is a block diagram of a data synchronization verification device provided in an embodiment of the present application;
[0021] FIG5 is a block diagram of an electronic device provided in an embodiment of the present application.
[0022] Implementation Methods of the Application
[0023] The following description of exemplary embodiments of the present application is made in conjunction with the accompanying drawings, including various details of the embodiments of the present application to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0024] In the absence of conflict, the various embodiments of the present application and the various features therein may be combined with each other.
[0025] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0026] The terms used herein are only used to describe specific embodiments and are not intended to limit this application. As used herein, the singular forms "a" and "the" are also intended to include the plural forms, unless the context clearly indicates otherwise. It will also be understood that when the terms "comprising" and / or "made of" are used in this specification, the presence of the features, wholes, steps, operations, elements and / or components is specified, but the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or groups thereof is not excluded. Similar words such as "connected" or "connected" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect.
[0027] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this application, and will not be interpreted as having an idealized or overly formal meaning unless expressly defined as such herein.
[0028] The embodiments of the present application provide a data synchronization verification method and device, an electronic device, and a computer-readable storage medium, which have the characteristics of short time consumption, less resource occupation, low delay, high verification efficiency, and reduced labor costs, thereby enabling accurate verification of data synchronization quality with fewer resources.
[0029] The data synchronization verification method according to the embodiment of the present application can be applied to the verification end, that is, the method is executed by an electronic device such as a terminal device or server corresponding to the verification end, and the terminal device can be a vehicle-mounted device, a user equipment (User Equipment, UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (Personal Digital Assistant, PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The server can be an independent physical server, a server cluster composed of multiple physical servers, or a cloud server capable of cloud computing. The method can be implemented by a processor calling a computer-readable program instruction stored in a memory. Alternatively, the method can be executed by a server.
[0030] Figure 1 is a flow chart of a data synchronization verification method provided by an embodiment of the present application. The method is applied to a verification terminal, which is used to implement each step of the data synchronization method. Referring to Figure 1, the method includes steps S11-S13, which are described in detail below.
[0031] In step S11, first characteristic information of first data and second characteristic information of second data are received, wherein the first data and the second data come from different data sources in a data synchronization process and include different batches of data.
[0032] In the embodiments of the present application, as an example, a process of data synchronization between a production end and a consumer end is used as an example. The data source of the first data is sent by the production end, and the data source of the second data can be the consumer end. The data synchronization process can be used to synchronize data from the production end to the consumer end. For example, the data synchronization process can include the execution of a data synchronization task to synchronize data from the production end to the consumer end. The data synchronization process can be a near-real-time synchronization process. For example, it can refer to splitting a periodically executed synchronization task into multiple sub-synchronization tasks, where the execution time interval of each sub-synchronization task is shorter than the time interval of the periodically executed synchronization task, and the sum of the time intervals of the sub-synchronization tasks is equal to the time interval of the periodically executed synchronization task. For example, a synchronization task executed once a day can be split into multiple sub-synchronization tasks with shorter time periods. For example, the short time period can be 30 minutes or 1 hour. This can not only shorten the completion time of data synchronization, but also facilitate the business side to quickly analyze the near-real-time synchronized data and save peak resources for synchronization tasks. However, in the near-real-time data synchronization scenario, high requirements are also placed on how to ensure the quality of data synchronization.
[0033] Exemplarily, the data synchronization process may refer to data synchronization between a database management system and a data lake, for example, data synchronization between a database management system and a data lake is implemented based on a computing engine. Exemplarily, an embodiment of the present application may be applied to a data synchronization application scenario based on Apache Spark (hereinafter referred to as Spark) to synchronize MySQL to write to Hudi. Among them, Spark is a fast and general computing engine designed for large-scale data processing, MySQL is a relational database management system, and Hudi is a scan-optimized data storage abstraction for analytical services. It can support changes to data sets within a delay of minutes, and also supports incremental processing of this data set by downstream systems. It can be used in offline, quasi-real-time application scenarios. When the embodiment of the present application is applied to the above-mentioned data synchronization scenario, by completing the synchronization verification of each batch in a relatively timely manner, it can ensure the synchronization quality of the data and ensure the stability of Hudi under large-scale use. It should be noted that the embodiment of the present application can be applied to any data synchronization scenario, and is not limited to the application scenario of quasi-real-time MySQL synchronization with Hudi described above.
[0034] In an embodiment of the present application, the production end may include a system for collecting and sending data in a data synchronization task. For example, the production end may provide a service for collecting MySQL and synchronize the collected first data to the consumer end. The consumer end may include a system for receiving and storing data in a data synchronization task. For example, the consumer end may provide a service for storing Hive, wherein Hive is a data warehouse tool that can be used to extract, transform, load, etc. data. Among them, data synchronization can be that the production end directly sends the collected first data to the consumer end for the consumer end to receive and store as second data, or the production end sends the collected first data to the intermediate end, and synchronizes the first data to the consumer end through the intermediate end. For example, a production end based on Canal can collect MySQL data and send the collected first data to an intermediate end based on Kafka. The intermediate end based on Kafka sends the data from the production end to a consumer end based on Spark, and the consumer end can store the received data in Hudi. Among them, Canal is a middleware that provides incremental data subscription and consumption based on database incremental log parsing. It can support the parsing of MySQL logical logs and can be used to process the obtained data after the parsing is completed. Kafka is a high-throughput distributed publish-subscribe messaging system that can be used to process action stream data.
[0035] It should be noted that the production end and the consumer end of the embodiment of this application can be any form of two ends for performing data synchronization tasks, and this application does not limit the data synchronization method between the production end and the consumer end. For ease of understanding, the following example uses the production end reading MySQL and synchronizing the read first data to the consumer end through the intermediate end Kafka, and the consumer end storing the received second data in Hudi.
[0036] In an embodiment of the present application, the first characteristic information of the first data is various types of information that can characterize the first data sent to the consumer end and is obtained by the production end within its own business scope during data synchronization, and can be used for synchronization verification. For example, the production end can periodically or non-periodically perform statistics based on various types of information of the currently read first data. For example, it can count the message generation of each synchronization table and determine the first generation time of each first data. The production end can send the first characteristic information of the first data obtained by statistics to the verification end at a fixed time or when the information of a synchronization table is obtained by statistics.
[0037] The first characteristic information may include at least one of the first synchronization type, the first data volume, the first primary key information, and the first generation time of each piece of first data. It should be understood that those skilled in the art may set the first characteristic information according to actual circumstances, as long as the first characteristic information can represent various types of information of the first data sent to the consumer. This application does not limit the method by which the producer determines and sends the first characteristic information, or the form and content of the first characteristic information.
[0038] In an embodiment of the present application, the production end may also print any log information related to the data transmission status, and the log information may be used to detect the transmission status of each piece of first data. When it is determined that the synchronization verification result of the current batch of data is verification failure, the log information may be used to perform an abnormal analysis of the missing data and locate the cause of the data synchronization error, for example, to locate the operation that caused the data synchronization error. Exemplarily, the log information may include the first primary key information of each piece of first data, any time point information related to the first data (for example, the first generation time of the first data, etc.), and various point information generated by the interaction with the intermediate end.
[0039] In an embodiment of the present application, the second characteristic information of the second data is various types of information obtained by the consumer end in data synchronization within its own business scope and can characterize the second data received from the production end, which can be used for synchronization verification. For example, the consumer end can periodically or non-periodically perform statistics based on various types of information of the received second data when the data lake is successfully written, determine the second characteristic information of the second data, and send the second characteristic information of the second data to the verification end. Among them, the second characteristic information may include at least one of the second synchronization type, second data volume, second primary key information, and second generation time of each second data. For any second data, its second generation time at the consumer end may be the first generation time of the second data at the production end. It should be understood that those skilled in the art can set the second characteristic information according to actual conditions. This application does not limit the way in which the consumer end determines and sends the second characteristic information, as well as the form and content of the second characteristic information.
[0040] In an embodiment of the present application, the consumer end may also print corresponding log information, which may be used to perform anomaly analysis on missing data and locate the cause of the data synchronization error. For example, the log information may include at least the second primary key information of each second data.
[0041] In an embodiment of the present application, not only can data synchronization be achieved, but the production end and the consumer end can also respectively count the first characteristic information of the first data and the second characteristic information of the second data in the synchronization process, and send them to the verification end. The verification end performs data aggregation based on the first characteristic information of the first data received from the production end and the second characteristic information of the second data received from the consumer end. For example, the data is aggregated according to batches, and synchronization verification is performed when the first characteristic information of the first data and the second characteristic information of the second data of the current batch are obtained through aggregation.
[0042] In step S12 , it is determined whether the first data of the current batch is consistent with the second data according to the first characteristic information of the first data of the current batch and the second characteristic information of the second data of the current batch.
[0043] The first data of the current batch is part of all the first data sent by the production end. Optionally, the first data sent by the production end can be divided into multiple batches based on information such as the production time and data volume of the first data. This allows data synchronization verification to be performed on the first data and second data of each batch. The second data of the current batch can be understood as the data of the first data of the current batch that is synchronized to the consumer end.
[0044] Based on the first characteristic information of the first data of the current batch and the second characteristic information of the second data of the current batch, it is possible to determine whether the first data of the current batch is consistent with the second data. For example, the first characteristic information is the data volume and primary key information of the first data, and the second characteristic information is the data volume and primary key information of the second data. Based on whether the data volume of the first data is consistent with the data volume of the second data, and whether the primary key information of the first data is consistent with the primary key information of the second data, it is determined whether the first data and the second data of the current batch are consistent.
[0045] In step S13 , when the first data and the second data of the current batch are consistent, it is determined that the result of the data synchronization verification of the current batch is verification passed.
[0046] If the first data and the second data of the current batch are consistent, the data synchronization verification result of the current batch can be determined to be passed. Conversely, if the first data and the second data of the current batch are inconsistent, the verification end can determine that the data synchronization verification result of the current batch is failed. Furthermore, missing data can be determined and anomaly analysis can be performed on the missing data.
[0047] According to the embodiments of the present application, it is possible to obtain the characteristic information of the first data sent by the production end and the characteristic information of the second data received by the consumer end during the data synchronization process between the production end and the consumer end, and perform data synchronization verification based on the characteristic information of the first data and the characteristic information of the second data. By determining whether the first data and the second data are consistent, it is verified whether the first data in the production end is completely synchronized to the second data in the consumer end. Since the characteristic information is part of the information in the data and can be used to represent important characteristics of the data, the complexity of the synchronization verification can be reduced and the resources occupied in the verification process can be reduced. In addition, the data synchronization verification process is verified according to data batches. The first data sent by the production end and the second data sent by the consumer end are verified in batches, which is conducive to reducing the amount of data processing for each verification and improving the verification efficiency.
[0048] According to the embodiments of the present application, the characteristic information of the first data and the second data of the current batch can be determined more timely by receiving the first characteristic information of the first data and the second characteristic information of the second data; synchronization verification is performed based on whether the characteristic information is consistent, and lightweight synchronization verification is performed according to the batch, without the need to perform heavyweight verification tasks at the level of a single data piece, which can reduce resource usage. Moreover, lightweight synchronization verification performed according to the batch can judge the data synchronization status in real time. Compared with the synchronization verification method of random sampling, the synchronization verification result obtained is more accurate and more convincing. According to the embodiments of the present application, it is possible to reduce the verification time consumption, reduce the verification delay, effectively improve the verification efficiency, and improve the accuracy of the verification results, thereby achieving accurate verification of data synchronization with fewer resources and in a more timely manner.
[0049] The following describes the data synchronization verification method according to an embodiment of the present application.
[0050] In some possible implementations, the verification end may perform synchronization verification based on the first characteristic information of the first data of the batch and the second characteristic information of the second data when it is determined that there is a batch without synchronization delay and the first characteristic information of the first data of the batch and the second characteristic information of the second data have been collected completely.
[0051] For example, the data generation speed of the production end fluctuates. For example, in certain time periods, the data generation speed of the production end is relatively low, and the consumer end can receive and store data in a timely manner, that is, the synchronization delay of the consumer end is relatively small. In certain time periods, the data generation speed of the production end is relatively high, and when the consumer end cannot receive and store a large amount of data in a timely manner, the synchronization delay of the consumer end is relatively large. For example, the production end generates a large amount of first data at 1 o'clock, and the consumer end may not be able to fully process it until 3 o'clock. When it is fully processed, it can be determined that there is a batch without synchronization delay, and before 3 o'clock, it can be determined that there is no batch without synchronization delay. As mentioned above, the first characteristic information of the first data and the second characteristic information of the second data are used for synchronization verification. When it is determined that there is a current batch without synchronization delay, and the first characteristic information of the first data of the current batch and the second characteristic information of the second data have been completely collected, synchronization verification can be performed based on the first characteristic information of the first data of the current batch and the second characteristic information of the second data.
[0052] In an embodiment of the present application, verification is performed when it is determined that there are batches without synchronization delays and the first characteristic information of the first data and the second characteristic information of the second data used for different verifications have been completely collected. This ensures that the data used for verification is accurate and complete, thereby ensuring the accuracy of the verification. It should be noted that some of the data read by the production end is the average for the entire day, some is uneven throughout the day, and some is generated at regular intervals. The embodiment of the present application does not waste a lot of resources to ensure low latency for all sub-synchronization tasks.
[0053] In an embodiment of the present application, batches can be divided according to time zones. In some optional application scenarios, they can also be divided according to business types. For example, the synchronized data includes data of multiple business types, and each business type can be a batch. As long as the data range included in each batch can be clearly defined, the present application does not limit the way the batches are divided.
[0054] In some possible implementations, batches can be divided according to time intervals. For example, batches can be divided according to hourly time intervals, where the lengths of time intervals for different batches can be the same or different. For example, they can be summarized once every hour, or the lengths of time intervals for different batches can be set to be different. This application does not limit the lengths of time intervals for batches. In this way, the verification frequency of synchronization verification can be flexibly adjusted by setting the lengths of time intervals for different batches according to the characteristics of the business scenario, which can effectively save verification resources.
[0055] In some possible implementations, the first characteristic information includes a first generation time for each piece of first data, and the second characteristic information includes a second generation time for each piece of second data. The first generation time of the first characteristic information and the second generation time of the second characteristic information of the same batch of second data fall within the same time interval.
[0056] As previously mentioned, for any piece of data, its second generation time at the consumer end can be the first generation time of the data at the producer end. When batches are divided by time intervals, the first generation time of the first characteristic information and the second generation time of the second characteristic information of the same batch fall within the same time interval.
[0057] In this way, the verification end can more easily and accurately determine the data range and data consumption of the same batch, thereby facilitating the verification end to determine whether the conditions for performing order-of-magnitude data verification are met. For example, based on the second generation time of the received second characteristic information, it can be determined whether the conditions for synchronous verification of the current batch of data are met. For example, based on the second generation time of the received second characteristic information, it can be determined whether there is a current batch without synchronization delay and whether the first characteristic information and second characteristic information of the current batch have been fully collected to achieve synchronous verification.
[0058] In some possible implementations, the method further includes: when the second generation time of the received second characteristic information exceeds the time interval corresponding to the current batch, determining that the first characteristic information and the second characteristic information of the current batch have been completely collected; and determining the first characteristic information of the first data and the second characteristic information of the second data of the current batch based on the first generation time and the second generation time.
[0059] For example, during the data synchronization process, the verification end has been receiving characteristic information sent from the production end and the consumption end, and may receive characteristic information of multiple data at a time. Assume that the time interval of the current batch is from 14:00 to 15:00, that is, the production time of the data at the production end is between 14:00 and 15:00. If the second generation time of the second characteristic information received by the verification end has exceeded 15:00, for example, the earliest production time received in the second characteristic information is 15:01, it means that there is no data at an earlier time point, and it can be determined that the first characteristic information and the second characteristic information of the current batch have been collected completely. Furthermore, based on the first generation time of each piece of first characteristic information received and the second generation time of each piece of second characteristic information received, the first characteristic information and the second characteristic information of the current batch can be summarized. For example, each piece of first characteristic information including the first generation time between 14:00 and 15:00 can be determined as the first characteristic information of the current batch, and each piece of second characteristic information including the second generation time between 14:00 and 15:00 can be determined as the second characteristic information of the current batch.
[0060] In some possible implementations, data from the same batch can be aggregated, updated, and stored in a single record. For example, the record may include multiple fields of information, such as batch information, the first data volume, synchronization type (including data insertion, data update, data deletion, etc.), primary key information, and first generation time of each first data item synchronized within the current batch on the production side; and the second data volume, synchronization type (including data insertion, data update, data deletion, etc.), primary key information, and second generation time of each second data item synchronized within the current batch on the consumer side.
[0061] In this way, when the verification end determines that the second generation time of the second characteristic information exceeds the time interval, it can determine that the first data of the production end of the current batch has been synchronized and received and stored at the consumer end, so that it can accurately judge that the first characteristic information and the second characteristic information of the current batch have been collected completely, and further determine the first characteristic information and the second characteristic information of the current batch for synchronization verification.
[0062] In some possible implementations, the second characteristic information of the second data is obtained by periodically or aperiodically aggregating the second characteristic information of the second data received by the consumer end, wherein the second generation time used to determine whether the first characteristic information and the second characteristic information of the current batch have been fully collected is earlier than other second generation times included in the second characteristic information. For example, the second characteristic information may include multiple second generation times, and if the earliest second generation time among the multiple second generation times has exceeded the time interval, it can be directly determined that the first characteristic information and the second characteristic information of the current batch have been fully collected.
[0063] In this way, when the earliest second generation time among multiple second generation times of the second characteristic information has exceeded the time interval, it can be ensured that the second characteristic information of subsequent second data including the second characteristic information no longer includes the data of the current batch, so that it can be quickly and accurately determined that the first data of the production end of the current batch has been synchronized and received and stored at the consumer end, reducing the probability of data omission and improving the accuracy of data verification.
[0064] In an embodiment of the present application, in step S12, whether the first data of the current batch is consistent with the second data is determined based on the first characteristic information of the first data of the current batch and the second characteristic information of the second data of the current batch. The specific method of determining the consistency of the first data and the second data can be based on the contents of the first characteristic information and the second characteristic information.
[0065] In the case that different synchronization types exist for data during the data synchronization process, the data consistency judgment may include a consistency judgment of the data amount and a consistency judgment of the synchronization type.
[0066] In some possible implementations, the first characteristic information includes a first synchronization type and a first data volume for each piece of first data, and the second characteristic information includes a second synchronization type and a second data volume for each piece of second data. Step S12 may include: determining whether the first data volume of the first data of the current batch is consistent with the second data volume of the second data; when the first data volume and the second data volume are consistent, determining whether the first synchronization type of each piece of first data is consistent with the second synchronization type of each piece of second data; when the first synchronization type of each piece of first data is consistent with the second synchronization type of each piece of second data, determining that the first data of the current batch is consistent with the second data.
[0067] That is to say, the first data volume of each first data and the second data volume of each second data can be compared to determine whether the first data volume of the first data of the current batch is consistent with the second data volume of the second data; if the data volumes are consistent, a synchronization type comparison can also be performed to compare whether the first synchronization type of each first data and the second synchronization type of each second data are consistent; if the synchronization types are consistent, it can be determined that the first data of the current batch is consistent with the second data; and then in step S13, it is determined that the data synchronization verification result of the current batch is verification passed.
[0068] The verification method based on the total data volume and synchronization type consistency judgment can reduce the resources occupied by synchronization verification and improve the accuracy of verification.
[0069] In some possible implementations, the first characteristic information includes a first synchronization type and a first data volume for each piece of first data, and the second characteristic information includes a second synchronization type and a second data volume for each piece of second data; the synchronization types of the data synchronization process include data insertion, data update, and data deletion. Step S12 may include: determining the first data volume of data insertion, the first data volume of data update, and the first data volume of data deletion based on the first synchronization type of the first data of the current batch; determining the second data volume of data insertion, the second data volume of data update, and the second data volume of data deletion based on the second synchronization type of the second data of the current batch; if the first data volume of data insertion is consistent with the second data volume of data insertion, the first data volume of data update is consistent with the second data volume of data update, and the first data volume of data deletion is consistent with the second data volume of data deletion, then it is determined that the first data and the second data of the current batch are consistent.
[0070] That is to say, according to the first synchronization type of each first data in the first characteristic information, the number of first data of type data insertion, the number of first data of type data update, and the number of first data of type data deletion can be determined respectively; similarly, according to the second synchronization type of each second data in the second characteristic information, the number of second data of type data insertion, the number of second data of type data update, and the number of second data of type data deletion can be determined respectively; then, the number of data of type data insertion, the number of data of type data update, and the number of data of type data deletion are compared respectively; if the number of data of each synchronization type is consistent, it can be determined that the first data of the current batch is consistent with the second data; then, in step S13, it is determined that the data synchronization verification result of the current batch is verification passed.
[0071] The accuracy of synchronization verification can be improved by using a verification method that determines the amount of data of multiple synchronization types separately. For example, when the synchronization types are not distinguished and all data amounts are compared for consistency, the total amount of data of each synchronization type may be consistent, but the amount of data of some synchronization types may be inconsistent. In this case, not distinguishing the synchronization types and comparing all data amounts for consistency may lead to misjudgment. The above consistency judgment method provided in the embodiment of the present application will give an accurate result of data amount inconsistency, thereby improving the accuracy of data amount consistency judgment.
[0072] In an embodiment of the present application, when the feature information includes primary key information, the consistency determination of the first data and the second data may also include a determination of the primary key information. The primary key information is the primary key of the data, which is one or more fields in the data that uniquely identifies a piece of data.
[0073] In some possible implementations, the first characteristic information further includes first primary key information for each piece of first data, and the second characteristic information further includes second primary key information for each piece of second data. Step S12 further includes: matching the first primary key information for each piece of first data with the second primary key information for each piece of second data; if a match is successful, determining that the first data and the second data in the current batch are consistent.
[0074] That is to say, the first primary key information of each first data can be matched with the second primary key information of each second data; if the first primary key information of each first data can be matched to the same second primary key information, the match is determined to be successful, and it can be determined that the first data and the second data of the current batch are consistent.
[0075] The accuracy of synchronization verification can be further improved through the primary key matching verification method.
[0076] In an embodiment of the present application, when the content of the characteristic information includes the generation time of the data, the consistency judgment of the first data and the second data may also include the judgment of the generation time.
[0077] In some possible implementations, the first characteristic information further includes a first generation time for each piece of first data, and the second characteristic information further includes a second generation time for each piece of second data. Step S12 further includes: matching the first generation time for each piece of first data with the second generation time for each piece of second data; if a match is successful, determining that the first data and the second data of the current batch are consistent.
[0078] That is, the first generation time of each piece of first data can be matched with the second generation time of each piece of second data. If the first generation time of each piece of first data can be matched to the same second generation time, the match is determined to be successful, and the first data and second data of the current batch can be determined to be consistent. Then, in step S13, the data synchronization verification result of the current batch is determined to be verified as passed.
[0079] By generating a time matching verification method, the accuracy of synchronization verification can be further improved.
[0080] The above describes the synchronization verification methods for judging the consistency of the total data volume and synchronization type, judging the number of data of multiple synchronization types, matching primary keys, and matching generation time. However, it should be understood that those skilled in the art may use any of the above-mentioned synchronization verification methods alone or in combination according to actual conditions, and this application does not impose any restrictions on this.
[0081] In some possible implementations, the method further includes: when the first data and the second data of the current batch are inconsistent, determining that the data synchronization verification result of the current batch is verification failure; determining the third primary key information of the missing data of the current batch based on the first primary key information of the first data of the current batch and the second primary key information of the second data of the current batch; and obtaining log information of the missing data based on the third primary key information of the missing data of the current batch, and the log information is used to perform anomaly analysis on the missing data.
[0082] For example, if the first data and the second data of the current batch are inconsistent, the verification end can determine that the data synchronization verification result of the current batch is verification failure. In the case of verification failure, the verification end can perform an alarm operation and conduct further analysis on the batch that failed the verification through a full comparison method across data sources. For example, the first primary key information of each first data in the current batch is matched with the second primary key information of each second data to determine the third primary key information of the missing data in the current batch. For example, a synchronization verification task can be started autonomously, which reads the MySQL and Hive tables of the batch that failed the verification and determines the third primary key information of the missing data in the batch.
[0083] In this way, there is no need to perform single-data-level comparison on all synchronized data. Instead, if the verification fails, single-data-level comparison is performed on the batch that failed the verification. This can effectively reduce the amount of data to be compared, ensure the accuracy of the synchronization verification, improve the comparison efficiency, and promptly determine the primary key information of the missing data in the batch.
[0084] Based on the third primary key information of the missing data in the current batch, the second data missing on the consumer side can be determined, and specifically which first data on the production side it corresponds to. Thus, by obtaining the log information of the missing data, the cause of the synchronization anomaly of the missing data can be determined.
[0085] In some possible implementations, when synchronizing first data from a production end to a consumer end via an intermediary, log information for the missing data at the sending end, consumer end, and intermediary end can be obtained based on the third primary key information of the missing data in the current batch, thereby assisting in locating the cause of data loss during the data synchronization process. The intermediary end is used to synchronize the first data from the production end to the consumer end, and the log information is used to perform anomaly analysis on the missing data.
[0086] Among them, a corresponding exception analysis algorithm can be designed to determine the error stage of the missing data based on the log information; the log information can also be sent to relevant personnel responsible for exception analysis (such as operation and maintenance personnel), and the relevant personnel will implement the exception analysis of the missing data. This application does not impose any restrictions on this.
[0087] For example, in a scenario where the first data from the production end is synchronized to the consumer end through the intermediate end, the error stage that causes missing data can be any stage from the production end reading data to the consumer end storing data, for example, it can include the stage from the production end reading data, the stage from the production end sending the first data to the intermediate end, the stage from the intermediate end processing data, the stage from the consumer end obtaining data from the intermediate end, and the stage from the consumer end storing data, etc. Log information can be used to record any data processing record at each end, for example, various operations on data. In the event of missing data, an exception analysis can be performed based on the various log information in this batch to determine the specific error stage that caused the missing data. For example, based on the log information of the missing data, determine at which stage the missing data is abnormal, and determine the operations related to the missing data in the specific error stage to locate the error cause of the missing data. This application does not limit the specific method of locating the error cause based on the log information.
[0088] According to the embodiments of the present application, during the data synchronization process, the relevant information used for lightweight synchronization verification is buried, collected, counted and alarmed at the verification end, so that the synchronization verification of the synchronized data can be carried out in a timely and efficient manner. As for inconsistent data volume and batches that have been determined to have failed the synchronization verification, the embodiments of the present application support heavyweight verification and comparison at the level of a single data item, and can assist in locating the root cause of data loss during the data synchronization process based on the primary key information of the missing data and the log information of the batch. In addition, the stored data can also be used as a source of report data for processing the synchronization task execution report and sending it to the user of the synchronization task.
[0089] Figure 2 is a schematic diagram of an application scenario of a data synchronization verification method provided in an embodiment of the present application. Figure 3 is a schematic diagram of an application scenario of a data synchronization verification method provided in an embodiment of the present application. For ease of understanding, the data synchronization verification process of an embodiment of the present application will be illustrated below in conjunction with Figures 2 and 3.
[0090] As shown in Figure 2, the producer reads first data (i.e., MySQL data) and appends relevant Hudi parameters. The intermediary Kafka server then synchronizes the read first data to the consumer server based on the intermediary. The intermediary Kafka server creates a data information queue topic for synchronizing data with the consumer server. The producer server collects first feature information for synchronization verification, such as the synchronization type and the first generation time of the first data. The consumer server parses the MySQL parameters and stores them in Hudi, consuming the received second data. Furthermore, the consumer server aggregates the synchronization type, second generation time, and location information of the second data. The intermediary Kafka server creates a feature information queue topic for synchronization verification, which sends the first feature information of the first data on the producer server and the second feature information of the second data on the consumer server to the verification server. The verification server analyzes the received information, namely the first feature information of the first data and the second feature information of the second data. If the lightweight synchronization verification conditions for a batch are met (data collection is complete), the verification server parses the relevant data, aggregates the first data on the producer server and the second data on the consumer server into a single record, and performs synchronization verification to determine whether the first and second data in the batch are consistent. When the first data and the second data are consistent, the data synchronization verification result of the batch is determined to be passed; when the first data and the second data are inconsistent, the data synchronization verification result of the batch is determined to be failed and an alarm is issued.
[0091] As shown in Figure 3, taking the example of a time interval where each hour is a batch, for example, 13:00 to 14:00 is a batch interval, and 14:00 to 15:00 is another batch interval. The verification end can determine the completion time of the current task synchronization queue and round up the completion time. The rounded-up result is the hourly hour, and the hourly hour is earlier than the completion time. For example, if the completion time is 15:03, the rounded-up result is 15:00. If the completion time is 14:59, the rounded-up result is 14:00. Based on the rounded-up hourly hourly number, the verification end can determine whether there are batches with no synchronization delays and complete the batch verification. In this application scenario, it can determine whether there are any hourly intervals without delays. For example, if the rounded-up result is 14:00, it can be determined that there may still be data in the time interval from 14:00 to 15:00 that has not completed data synchronization. Since there are currently no batches with no synchronization delays, that is, no hourly intervals without delays, synchronization verification is not performed and the test is skipped. For example, if the value is rounded up to 15 o'clock, it can be determined that the data in the time interval of the batch from 14 o'clock to 15 o'clock has completed data synchronization, and there is currently a batch without synchronization delay, that is, there is an hour interval without delay. The verification end performs the synchronization verification described above based on the hour interval without delay. If the data is consistent, the data synchronization verification result of the batch is determined to be passed, no alarm is issued, and the test ends. If the data is inconsistent, the data synchronization verification result of the batch is determined to be failed, an alarm is issued, and a heavyweight synchronization verification task is further initiated to determine the primary key information of the missing data in the batch. The corresponding log information can also be stored for further investigation of the cause of the data synchronization error.
[0092] It is understood that the various method embodiments mentioned in this application can be combined with each other to form combined embodiments without violating the principle logic. Due to space limitations, this application will not go into details. Those skilled in the art will understand that in the above-mentioned methods of specific implementation, the specific execution order of each step should be determined by its function and possible internal logic.
[0093] In addition, the present application also provides a data synchronization verification device, a data synchronization system, an electronic device, and a computer-readable storage medium, all of which can be used to implement any data synchronization verification method provided by the present application. The corresponding technical solutions and descriptions are referred to the corresponding records in the method section and will not be repeated here.
[0094] FIG4 is a block diagram of a data synchronization verification device provided in an embodiment of the present application.
[0095] 4 , an embodiment of the present application provides a data synchronization verification device, which includes a receiving module 41, a consistency determination module 42 and a verification result determination module 43. The receiving module 41 is configured to receive first characteristic information of first data and second characteristic information of second data, wherein the first data and the second data respectively come from different data sources in the data synchronization process and respectively include different batches of data. The consistency determination module 42 is configured to determine whether the first data of the current batch is consistent with the second data based on the first characteristic information of the first data of the current batch and the second characteristic information of the second data of the current batch. The verification result determination module 43 is configured to determine that the data synchronization verification result of the current batch is verification passed when the first data of the current batch is consistent with the second data.
[0096] In some possible implementations, the first characteristic information includes a first synchronization type and a first data volume for each piece of first data, and the second characteristic information includes a second synchronization type and a second data volume for each piece of second data; the consistency determination module 42 is further configured to: determine whether the first data volume of the first data of the current batch is consistent with the second data volume of the second data; when the first data volume and the second data volume are consistent, determine whether the first synchronization type of each piece of first data is consistent with the second synchronization type of each piece of second data; when the first synchronization type of each piece of first data is consistent with the second synchronization type of each piece of second data, determine that the first data of the current batch is consistent with the second data.
[0097] In some possible implementations, the first characteristic information includes a first synchronization type and a first data amount for each first data item, the second characteristic information includes a second synchronization type and a second data amount for each second data item, and the synchronization types of the data synchronization process include data insertion, data update, and data deletion; the consistency determination module 42 is further configured to: determine the first data amount of data insertion, the first data amount of data update, and the first data amount of data deletion according to the first synchronization type of the first data of the current batch; determine the second data amount of data insertion, the second data amount of data update, and the second data amount of data deletion according to the second synchronization type of the second data of the current batch; if the first data amount of data insertion is consistent with the second data amount of data insertion, the first data amount of data update is consistent with the second data amount of data update, and the first data amount of data deletion is consistent with the second data amount of data deletion, then it is determined that the first data and the second data of the current batch are consistent.
[0098] In some possible implementations, the first characteristic information also includes first primary key information for each piece of first data, and the second characteristic information also includes second primary key information for each piece of second data, where the primary key information is used to uniquely identify a piece of data; the consistency determination module 42 is also configured to: match the first primary key information for each piece of first data with the second primary key information for each piece of second data; if the match is successful, it is determined that the first data of the current batch is consistent with the second data.
[0099] In some possible implementations, the first characteristic information also includes a first generation time of each piece of first data, and the second characteristic information also includes a second generation time of each piece of second data; the consistency determination module 42 is further configured to: match the first generation time of each piece of first data with the second generation time of each piece of second data; if the match is successful, it is determined that the first data of the current batch is consistent with the second data.
[0100] In some possible implementations, the first characteristic information includes a first generation time of each piece of first data, and the second characteristic information also includes a second generation time of each piece of second data; the device further includes a collection determination module and a characteristic information determination module. The collection determination module is configured to determine that the first and second characteristic information of the current batch have been completely collected if the second generation time of the received second characteristic information exceeds a time interval corresponding to the current batch; and the characteristic information determination module is configured to determine, if the first and second characteristic information of the current batch have been completely collected, the first characteristic information of the first data and the second characteristic information of the current batch based on the first generation time and the second generation time.
[0101] In some possible implementations, the device further includes a verification result determination module, a primary key determination module, and a log acquisition module. The verification result determination module is configured to determine that the data synchronization verification result of the current batch is verification failure when the first data and the second data of the current batch are inconsistent; the primary key determination module is configured to determine the third primary key information of the missing data in the current batch based on the first primary key information of the first data of the current batch and the second primary key information of the second data of the current batch; and the log acquisition module is configured to obtain log information of the missing data based on the third primary key information of the missing data in the current batch, and the log information is used to perform anomaly analysis on the missing data.
[0102] FIG5 is a block diagram of an electronic device provided in an embodiment of the present application.
[0103] 5 , an embodiment of the present application provides an electronic device, comprising at least one processor 701, at least one memory 702, and one or more I / O interfaces 703. The one or more I / O interfaces 703 are connected between the processor 701 and the memory 702. The memory 702 stores one or more computer programs executable by the at least one processor 701, and the one or more computer programs are executed by the at least one processor 701 to enable the at least one processor 701 to perform the above-described data synchronization verification method.
[0104] The present application also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the above-mentioned data synchronization verification method when executed by a processor. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.
[0105] An embodiment of the present application also provides a computer program product, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above-mentioned data synchronization verification method.
[0106] It will be understood by those skilled in the art that all or some of the steps, systems, and functional modules / units in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. In a hardware implementation, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or may be implemented as hardware, or may be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable storage medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium).
[0107] As is well known to those skilled in the art, the term computer storage media includes volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information (such as computer-readable program instructions, data structures, program modules or other data). Computer storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technology, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those skilled in the art, communication media typically contains computer-readable program instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0108] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.
[0109] The computer program instructions for performing the operation of the present application can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data or source code or object code written in any combination of one or more programming languages, wherein the programming language includes object-oriented programming languages such as Smalltalk, C++, and conventional procedural programming languages such as "C" language or similar programming languages. Computer-readable program instructions can be executed completely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or executed completely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer by any type of network including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (such as by using an Internet service provider to connect to the Internet). In certain embodiments, by utilizing the state information of computer-readable program instructions to personalize electronic circuits, such as programmable logic circuits, field programmable gate arrays (FPGAs) or programmable logic arrays (PLAs), the electronic circuits can execute computer-readable program instructions, thereby realizing various aspects of the present application.
[0110] The computer program product described herein may be implemented in hardware, software, or a combination thereof. In one embodiment, the computer program product is implemented as a computer storage medium. In another embodiment, the computer program product is implemented as a software product, such as a software development kit (SDK).
[0111] Various aspects of the present application are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0112] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0113] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0114] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the system, method and computer program product according to multiple embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction includes one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a special hardware-based system that performs the function or action of the specification, or can be implemented by a combination of special hardware and computer instructions.
[0115] Example embodiments have been disclosed herein, and although specific terms are employed, they are used and should be interpreted only in a general illustrative sense and not for purposes of limitation. In some instances, it will be apparent to those skilled in the art that, unless otherwise expressly indicated, features, characteristics, and / or elements described in conjunction with a particular embodiment may be used alone or in combination with features, characteristics, and / or elements described in conjunction with other embodiments. Therefore, it will be understood by those skilled in the art that various changes in form and detail may be made without departing from the scope of the present application as set forth in the appended claims.
Claims
1. A data synchronization verification method, the method comprising: Receiving first characteristic information of first data and second characteristic information of second data, wherein the first data and the second data come from different data sources in a data synchronization process and include different batches of data; Determining whether the first data of the current batch is consistent with the second data according to the first characteristic information of the first data of the current batch and the second characteristic information of the second data of the current batch; When the first data and the second data of the current batch are consistent, it is determined that the result of the data synchronization verification of the current batch is verification passed.
2. The method according to claim 1, wherein: The first characteristic information includes a first synchronization type and a first data volume of each first data, and the second characteristic information includes a second synchronization type and a second data volume of each second data; The determining, based on the first characteristic information of the first data of the current batch and the second characteristic information of the second data of the current batch, whether the first data of the current batch is consistent with the second data includes: Determine whether a first data amount of the first data of the current batch is consistent with a second data amount of the second data; In a case where the first data amount and the second data amount are consistent, determining whether a first synchronization type of each piece of first data is consistent with a second synchronization type of each piece of second data; In a case where the first synchronization type of each piece of the first data is consistent with the second synchronization type of each piece of the second data, it is determined that the first data of the current batch is consistent with the second data.
3. The method according to claim 1, wherein: The first characteristic information includes a first synchronization type and a first data volume of each first data, the second characteristic information includes a second synchronization type and a second data volume of each second data, and the synchronization type of the data synchronization process includes data insertion, data update and data deletion; The determining, based on the first characteristic information of the first data of the current batch and the second characteristic information of the second data of the current batch, whether the first data of the current batch is consistent with the second data includes: Determine, according to the first synchronization type of the first data of the current batch, the first data quantity of data insertion, the first data quantity of data update, and the first data quantity of data deletion; Determine, according to the second synchronization type of the second data of the current batch, the amount of second data for data insertion, the amount of second data for data update, and the amount of second data for data deletion; If the first data quantity of the data inserted is consistent with the second data quantity of the data inserted, the first data quantity of the data updated is consistent with the second data quantity of the data updated, and the first data quantity of the data deleted is consistent with the second data quantity of the data deleted, then it is determined that the first data and the second data of the current batch are consistent.
4. The method according to claim 2 or 3, wherein: The first characteristic information further includes first primary key information of each piece of first data, and the second characteristic information further includes second primary key information of each piece of second data, where the primary key information is used to uniquely identify a piece of data; Determining that the first data and the second data of the current batch are consistent includes: Matching the first primary key information of each piece of the first data with the second primary key information of each piece of the second data; If the match is successful, it is determined that the first data of the current batch is consistent with the second data.
5. The method according to claim 2 or 3, wherein: The first characteristic information further includes a first generation time of each piece of first data, and the second characteristic information further includes a second generation time of each piece of second data; Determining that the first data and the second data of the current batch are consistent includes: Matching a first generation time of each piece of the first data with a second generation time of each piece of the second data; If the match is successful, it is determined that the first data of the current batch is consistent with the second data.
6. The method according to claim 1, wherein: The first characteristic information includes a first generation time of each piece of first data, and the second characteristic information also includes a second generation time of each piece of second data; The method further comprises: When the second generation time of the received second characteristic information exceeds the time interval corresponding to the current batch, it is determined that the first characteristic information and the second characteristic information of the current batch have been completely collected; According to the first generation time and the second generation time, first characteristic information of the first data and second characteristic information of the second data of the current batch are determined.
7. The method according to claim 1, further comprising: When the first data and the second data of the current batch are inconsistent, determining that the data synchronization verification result of the current batch is verification failure; Determine the third primary key information of the missing data in the current batch according to the first primary key information of the first data in the current batch and the second primary key information of the second data in the current batch; According to the third primary key information of the missing data in the current batch, log information of the missing data is obtained, and the log information is used to perform abnormal analysis on the missing data.
8. A data synchronization verification device, the device comprising: A receiving module, configured to receive first characteristic information of first data and second characteristic information of second data, wherein the first data and the second data come from different data sources in a data synchronization process and include different batches of data; a consistency determination module, configured to determine whether the first data of the current batch is consistent with the second data according to the first characteristic information of the first data of the current batch and the second characteristic information of the second data of the current batch; The verification result determination module is configured to determine that the data synchronization verification result of the current batch is verification passed when the first data and the second data of the current batch are consistent.
9. The device according to claim 8, wherein: The first characteristic information includes a first synchronization type and a first data volume of each first data, and the second characteristic information includes a second synchronization type and a second data volume of each second data; The consistency determination module is further configured to: Determine whether a first data amount of the first data of the current batch is consistent with a second data amount of the second data; In a case where the first data amount and the second data amount are consistent, determining whether a first synchronization type of each piece of first data is consistent with a second synchronization type of each piece of second data; In a case where the first synchronization type of each piece of the first data is consistent with the second synchronization type of each piece of the second data, it is determined that the first data of the current batch is consistent with the second data.
10. The device according to claim 8, wherein: The first characteristic information includes a first synchronization type and a first data volume of each first data, the second characteristic information includes a second synchronization type and a second data volume of each second data, and the synchronization type of the data synchronization process includes data insertion, data update and data deletion; The consistency determination module is further configured to: Determine, according to the first synchronization type of the first data of the current batch, the first data quantity of data insertion, the first data quantity of data update, and the first data quantity of data deletion; Determine, according to the second synchronization type of the second data of the current batch, the amount of second data for data insertion, the amount of second data for data update, and the amount of second data for data deletion; If the first data quantity of the data inserted is consistent with the second data quantity of the data inserted, the first data quantity of the data updated is consistent with the second data quantity of the data updated, and the first data quantity of the data deleted is consistent with the second data quantity of the data deleted, then it is determined that the first data and the second data of the current batch are consistent.
11. The device according to claim 9 or 10, wherein: The first characteristic information further includes first primary key information of each piece of first data, and the second characteristic information further includes second primary key information of each piece of second data, where the primary key information is used to uniquely identify a piece of data; The consistency determination module is further configured to: Matching the first primary key information of each piece of the first data with the second primary key information of each piece of the second data; If the match is successful, it is determined that the first data of the current batch is consistent with the second data.
12. The device according to claim 9 or 10, wherein: The first characteristic information further includes a first generation time of each piece of first data, and the second characteristic information further includes a second generation time of each piece of second data; The consistency determination module is further configured to: Matching a first generation time of each piece of the first data with a second generation time of each piece of the second data; If the match is successful, it is determined that the first data of the current batch is consistent with the second data.
13. The device according to claim 8, wherein: The first characteristic information includes a first generation time of each piece of first data, and the second characteristic information also includes a second generation time of each piece of second data; The device also includes: a collection determination module, configured to determine that the first characteristic information and the second characteristic information of the current batch have been completely collected when the second generation time of the received second characteristic information exceeds the time interval corresponding to the current batch; The characteristic information determination module is configured to determine the first characteristic information of the first data and the second characteristic information of the second data of the current batch according to the first generation time and the second generation time when the first characteristic information and the second characteristic information of the current batch have been completely collected.
14. The device according to claim 8, further comprising: A verification result determination module, configured to determine that the data synchronization verification result of the current batch is verification failure when the first data and the second data of the current batch are inconsistent; A primary key determination module, configured to determine third primary key information of missing data of the current batch according to first primary key information of first data of the current batch and second primary key information of second data of the current batch; The log acquisition module is configured to acquire log information of the missing data according to the third primary key information of the missing data in the current batch, wherein the log information is used to perform abnormal analysis on the missing data.
15. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores one or more computer programs that can be executed by the at least one processor, and the one or more computer programs are executed by the at least one processor so that the at least one processor can perform the data synchronization verification method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Data verification method and device, computing equipment and storage medium
CN113626416A
Heterogeneous data source-oriented data consistency verification method and device
CN116150175A
Data verification method and device, computer equipment and storage medium
CN117093409A
Data synchronization verification method and device, electronic equipment and storage medium
CN117951144A
Data warehouse model validation
US20170293641A1
Cited By
Self-adaptive multi-stage filtering and accurate decoding data consistency checking system and self-adaptive multi-stage filtering and accurate decoding data consistency checking method
CN120723783A