Data processing method and device, storage medium and electronic equipment
By acquiring incremental data from the old and new business systems and performing type-based comparisons to form a differential data set, the problem of low accuracy in comparing the old and new systems is resolved, enabling more efficient data anomaly judgment.
Patent Information
- Application Number
- CN202510827305.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-09-26
AI Technical Summary
In the existing technology, the accuracy of data comparison between the new and old business systems is low, resulting in inaccurate judgment of data anomalies in the new system.
By obtaining incremental business data from the old and new business systems at multiple times, the data types belonging to the addition, modification, and deletion operations are compared to form a differential data set, and based on this differential data set, it is determined whether there are any anomalies in the new system.
The accuracy of data comparison between the old and new business systems has been improved, and it can more accurately determine whether there are data anomalies in the new system, avoiding misjudgments caused by differences in system processing time and the complexity of business logic.
Smart Images

Figure CN120705161A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and in particular to a data processing method, device, storage medium and electronic equipment. Background Art
[0002] With the development of the aviation business, to better serve target passenger groups and proactively respond to market changes, aviation application systems require architectural upgrades to meet expanding demand. Migrating from legacy systems to new ones presents numerous challenges. To avoid impacting downstream systems, the output of the old and new systems must be compared. Currently, relevant technologies rely on a consistency comparison of the full data from the old and new systems at a specific moment to determine whether the new system contains data anomalies, resulting in low accuracy.
[0003] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention
[0004] Embodiments of the present invention provide a data processing method, device, storage medium, and electronic device to at least solve the technical problem in the related art of low judgment accuracy when comparing business data of new and old business systems to determine whether there are data anomalies in the new system.
[0005] According to one aspect of an embodiment of the present invention, a data processing method is provided, including: respectively acquiring incremental business data of a first business system and a second business system at multiple moments to obtain a first business data set of the first business system and a second business data set of the second business system, wherein the second business system is a system to replace the first business system in processing business; comparing the incremental business data of the first type in the first business data set and the second business data set to obtain a first difference data set at each moment, wherein the first type refers to a data type associated with a new operation or a modification operation; comparing the incremental business data of the second type in the first business data set and the second business data set to obtain a second difference data set at each moment, wherein the second type refers to a data type associated with a deletion operation; judging whether there is an anomaly in the data in the second business system based on the first difference data set and the second difference data set to obtain target information.
[0006] Furthermore, the data processing method also includes: determining the incremental business data of the first type in the first business data set as the first business data, and determining the incremental business data of the first type in the second business data set as the second business data; for the first moment among multiple moments, determining the first difference data and the second difference data at the moment based on the first business data at the moment and the second business data at the moment, wherein the first difference data belongs to the first business data and the second difference data belongs to the second business data; for the Nth moment among multiple moments, determining the first difference data and the second difference data at the moment based on the first business data and the second business data at the moment, and the first difference data and the second difference data at the N-1th moment, wherein N is greater than 1; determining the first difference data set for each moment based on the first difference data and the second difference data at each moment.
[0007] Furthermore, the data processing method also includes: determining the primary key of the first business data at that moment as the first target primary key, deleting the data corresponding to the first target primary key in the first difference data at the N-1th moment to obtain the updated first difference data, and determining the first business data at that moment and the updated first difference data as the data to be compared at that moment; deleting the identical data between the data to be compared at that moment and the second difference data at the N-1th moment to obtain the updated data to be compared and the updated second difference data; deleting the identical data between the updated data to be compared and the second business data at that moment to obtain the updated data to be compared again and the updated second business data; determining the updated data to be compared again as the first difference data at that moment, and determining the updated second business data and the updated second difference data as the second difference data at that moment.
[0008] Furthermore, the data processing method also includes: determining the incremental business data belonging to the second type in the first business data set as third business data, and determining the incremental business data belonging to the second type in the second business data set as fourth business data; for the first moment among multiple moments, determining the third difference data and fourth difference data at the moment based on the third business data at the moment and the fourth business data at the moment, wherein the first difference data belongs to the third business data and the second difference data belongs to the fourth business data; for the Nth moment among multiple moments, determining the third difference data and fourth difference data at the moment based on the third business data and the fourth business data at the moment, and the first difference data, second difference data, third difference data and fourth difference data at the N-1th moment; determining the second difference data set for each moment based on the third difference data and fourth difference data at each moment.
[0009] Furthermore, the data processing method also includes: determining the primary key of the third business data at that moment as the second target primary key, and when there is no data corresponding to the second target primary key in the first difference data at the N-1th moment and the fourth difference data at the N-1th moment, determining the third business data as the third difference data at that moment; determining the primary key of the fourth business data at that moment as the third target primary key, and when there is no data corresponding to the third target primary key in the second difference data at the N-1th moment and the third difference data at the N-1th moment, determining the fourth business data as the fourth difference data at that moment.
[0010] Furthermore, the data processing method also includes: for the Mth moment among multiple moments, recording the first primary key of the moment to the first storage area, wherein the first primary key includes: a primary key for which first difference data exists but second difference data does not exist at the moment, a primary key for which third difference data exists but fourth difference data does not exist at the moment, and M is a positive integer; for the M+1th moment among multiple moments, if the first difference data or the third difference data of the first primary key of the Mth moment in the first storage area is not deleted, and the second business data and the fourth business data identical to the first primary key do not exist at the M+1th moment, then it is determined that the target information represents that there is an abnormality in the data in the second business system.
[0011] Furthermore, the data processing method also includes: for the Mth moment among multiple moments, recording the second primary key of the moment to the second storage area, wherein the second primary key includes: the primary key for which the second difference data exists and the first difference data does not exist at the moment, and the primary key for which the fourth difference data exists and the third difference data does not exist at the moment, and M is a positive integer; for the M+1th moment among the multiple moments, if the second difference data or the fourth difference data of the second primary key of the Mth moment in the second storage area is not deleted, and the first business data and the third business data identical to the second primary key do not exist at the M+1th moment, then it is determined that the target information represents that there is an abnormality in the data in the second business system.
[0012] Furthermore, the data processing method also includes: for the Mth moment among multiple moments, recording the third primary key of the moment to the third storage area, wherein the third primary key includes: the primary key of one of the first difference data and the third difference data exists at the moment, and the primary key of one of the second difference data and the fourth difference data exists, and M is a positive integer; if there is no incremental business data corresponding to the third primary key in the third storage area at the M+1th moment, it is determined that the target information represents that there is an abnormality in the data in the second business system.
[0013] According to another aspect of an embodiment of the present invention, a data processing device is also provided, including: an acquisition module, used to respectively acquire incremental business data of a first business system and a second business system at multiple times, to obtain a first business data set of the first business system and a second business data set of the second business system, wherein the second business system is a system to replace the first business system for processing business; a first comparison module, used to compare the incremental business data of the first type in the first business data set and the second business data set, to obtain a first difference data set at each moment, wherein the first type refers to a data type associated with a new operation or a modification operation; a second comparison module, used to compare the incremental business data of the second type in the first business data set and the second business data set, to obtain a second difference data set at each moment, wherein the second type refers to a data type associated with a deletion operation; a processing module, used to determine whether there is an abnormality in the data in the second business system based on the first difference data set and the second difference data set, to obtain target information.
[0014] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is configured to execute the above-mentioned data processing method when running.
[0015] According to another aspect of an embodiment of the present invention, an electronic device is also provided, which includes one or more processors; a memory for storing one or more programs, which enables the one or more processors to run the programs when the one or more programs are executed by the one or more processors, wherein the programs are configured to execute the above-mentioned data processing method when running.
[0016] According to another aspect of an embodiment of the present invention, a computer program product is provided, including a computer program / instruction, which implements the above-mentioned data processing method when the computer program / instruction is executed by a processor.
[0017] In an embodiment of the present invention, a method is adopted to determine whether the new system has data anomalies based on the incremental business data of the new and old systems at multiple times. By respectively obtaining the incremental business data of the first business system and the second business system at multiple times, a first business data set of the first business system and a second business data set of the second business system are obtained. Then, the incremental business data of the first type in the first business data set and the second business data set are compared to obtain a first difference data set at each moment. Then, the incremental business data of the second type in the first business data set and the second business data set are compared to obtain a second difference data set at each moment. Thus, based on the first difference data set and the second difference data set, it is determined whether the data in the second business system has an anomaly and the target information is obtained. Among them, the second business system is the system to replace the first business system for processing business. The first type refers to the data type associated with the characterization of the new operation or the modification operation, and the second type refers to the data type associated with the characterization of the deletion operation.
[0018] In the above process, by comparing the incremental business data belonging to the first type in the first business data set and the second business data set, and comparing the incremental business data belonging to the second type, and performing data anomaly judgment based on the first difference data set and the second difference data set obtained by the comparison, on the one hand, it is possible to determine whether there is an anomaly in the new system data by combining the difference information of the incremental business data of the new and old systems at multiple times, and to consider the data differences between the new and old systems caused by the system processing time, avoiding focusing only on the static consistency of the data, thereby improving the accuracy of judgment. On the other hand, it is possible to distinguish and compare the incremental business data in the new and old systems based on the different data types, thereby further improving the accuracy of judgment.
[0019] It can be seen from this that the solution provided by this application achieves the purpose of judging whether there are data anomalies in the new system based on the incremental business data of the new and old systems at multiple times, thereby achieving the technical effect of improving the accuracy of judging whether there are data anomalies in the new system, and further solving the technical problem of low judgment accuracy when comparing the business data of the new and old business systems in related technologies to judge whether there are data anomalies in the new system. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0021] Figure 1 It is a working diagram of a freight rate publishing system in the related art;
[0022] Figure 2 is a schematic diagram of optional incremental service data according to an embodiment of the present invention;
[0023] Figure 3 This is a flow chart of an optional data processing method according to an embodiment of the present invention. Figure 1 ;
[0024] Figure 4 is a schematic diagram of an optional comparison of incremental business data according to an embodiment of the present invention;
[0025] Figure 5 This is a flow chart of an optional data processing method according to an embodiment of the present invention. Figure 2 ;
[0026] Figure 6 is a schematic diagram of an optional method of comparing data using a differential algorithm according to an embodiment of the present invention;
[0027] Figure 7 is a schematic diagram of an optional data processing device according to an embodiment of the present invention;
[0028] Figure 8 is a schematic diagram of an optional electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0029] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0030] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0031] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0032] Example 1
[0033] According to an embodiment of the present invention, an embodiment of a data processing method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0034] With the development of aviation business, in order to better serve the target passenger groups and actively respond to market changes, the application systems in the aviation field need to be upgraded in architecture to meet the ever-expanding needs.
[0035] For example, the current airfare products (i.e., air ticket prices) in the aviation field involve systems such as Figure 1 As shown, Figure 1 This is a working diagram of the fare publishing system in the related art. After the information of the new or modified fare products of the airline is entered into the fare publishing system, the fare publishing system will process it into a data file at a short time interval, marking whether the data can currently be used for calculation and application in the market. The data file (i.e. Figure 1 The incremental data files and full data files in the data are pushed to the downstream computing and search systems, and the data files are stored in the freight rate publishing database. Figure 1 As shown, the operations on freight rate products entered into the freight rate publishing system include but are not limited to the following: 1. Creating a new freight rate product; 2. Modifying an existing freight rate product; 3. Terminate the current existing freight rate product; 4. Deleting an existing freight rate product.
[0036] Migrating from an old system to a new one involves many issues. To avoid impacting downstream systems, it's necessary to compare the outputs of the old and new systems. This application primarily involves comparing the files generated by the freight rate publishing system (i.e., the files output to downstream systems). This process presents the following difficulties:
[0037] Difficulty 1: The complexity of the file content itself in the case of business continuity. This part of the complexity stems from the continuity of data push. The operation of a piece of data at a certain point in time may be recorded in different push files in the process of generating data files due to the different time of entering the old and new systems or the different time of system processing. For example, the old system may record the operation of a certain data and enter the database at 14:49:59.008. Then, during the push process at 14:50, the operation can be pushed to the downstream system; the new system may record the operation time as 14:50:00.096. Then, during the push process at 14:50, the data operation will not be pushed, and the operation will be recorded in the next data push file. The non-synchronous nature of push operations may also cause this problem. Therefore, file comparison tools cannot simply draw conclusions based on whether the content of the file push is consistent.
[0038] Difficulty 2: Complexity of business push logic. The current push structure of additional service business data is as follows: Figure 2 , Figure 2This is a schematic diagram of an optional incremental business data according to an embodiment of the present invention. The incremental business data includes the data content of the data being operated (such as the data content of the data with table number 012 and the data logical primary key being "XXXXYY") and the type of operation on the data content (such as adding I, modifying C, deleting D). Here, the operations of deletion and modification are both for the data that have been pushed before, and the expiration operation is regarded as a modification of the effective time of the data. Therefore, the push process of the two systems for the data may be completely inconsistent in content but both are correct in terms of business. For example, for a certain piece of data, such as table number 012 and the data logical primary key being "XXXXYY", the user's operations on the data content are: adding operation L1, modifying operation L2, modifying operation L3, modifying operation L4, and deleting operation L5. At this time, the following normal situations may exist: (1) At time T1, the old system pushes the data file and recognizes the data adding operation L1. The new system has not yet recognized the addition of the data, and the push for this piece of data is empty. (2) At T2, the old system recognizes that the data has been modified to L3 and pushes the modification operation L3. The new system pushes the data file and recognizes that the data has been newly added L2 (at this time, the data has not been pushed to the new system before, and it is newly added to the downstream systems). (3) At T3, the old system recognizes that the data has been modified to L4 and pushes the modification operation L4. The new system recognizes that the data has been deleted and pushes the deleted data L2; (4) At T4, the old system recognizes that the data has been deleted and pushes the deleted data L4. The above process is logically reasonable for both the old and new systems to push, and neither should be considered an abnormal situation. However, the content pushed by the two systems is always inconsistent. For example, the old system may recognize the addition of a piece of data at T1 and the deletion of the same data at T2. For the new system, the data was deleted before it was pushed, so the pushed file has no record of the data at all, which is also reasonable. Therefore, the file comparison tool needs to be compatible with the rationality of the above business and identify abnormal data on this basis.
[0039] At present, relevant technologies conduct consistency comparison based on the full amount of data of the new and old systems at a specific moment to determine whether there are data anomalies in the new system, which leads to the problem of low judgment accuracy.
[0040] In the above technical background, this application provides Figure 3 How to handle the fault shown. Figure 3 This is a flow chart of an optional data processing method according to an embodiment of the present invention. Figure 1 .like Figure 3 As shown, the method includes:
[0041] Step S301, respectively obtain incremental business data of the first business system and the second business system at multiple times, to obtain a first business data set of the first business system and a second business data set of the second business system, wherein the second business system is the system to replace the first business system in processing business.
[0042] Optionally, electronic devices, application systems, servers and other devices may be used as the execution subject of the present application. In this embodiment, the target processing system is used as the execution subject to execute the above data processing method.
[0043] Optionally, the first business system can be understood as an existing, mature business system, which can be called an old system. The second business system can be understood as a newly developed or upgraded system intended to replace the first business system, which can be called a new system.
[0044] During the architecture upgrade, data input into the old system is synchronously replicated and diverted to the new system. The old and new systems each process the input data and output corresponding processing results (i.e., the incremental business data described above). During this process, the processing results output by the old system are used for actual applications and sent to downstream systems for business processing. The processing results output by the new system are compared with those of the old system to determine whether there are any problems with the new system in actual applications. Once the new system is finally confirmed to be operationally stable, it is used for actual business processing and the old system is no longer used.
[0045] In an optional embodiment, the first business system and the second business system may be a freight rate publishing system, which may be used to receive operations on freight rates (such as adding new freight rates, modifying freight rates, deleting freight rates, etc.), and then output corresponding freight rate incremental data, which is also incremental business data.
[0046] Optionally, the above-mentioned multiple moments refer to the time points when the first business system and the second business system push incremental business data. For example, the first business system and the second business system each push incremental business data every 15 minutes, and the push time points of the first business system and the second business system are the same.
[0047] Step S302 , compare the incremental business data of the first type in the first business data set and the second business data set to obtain a first difference data set at each moment, wherein the first type refers to a data type associated with a new operation or a modification operation.
[0048] Optional, such as Figure 2As shown, the first type of incremental business data may refer to data with the fourth digit identified as "I" or "C". The meaning of the first type of incremental business data may be "adding data content A to the primary key 112 in the 012 table" or "modifying the data at the primary key 100 in the 012 table to data content B".
[0049] Optionally, at each moment, the target processing system may compare the newly added and modified data at that moment with the first difference data set at the previous moment to find out inconsistent records and form the first difference data set at that moment.
[0050] Step S303 : Compare the incremental business data of the second type in the first business data set and the second business data set to obtain second difference data sets at each moment, wherein the second type refers to a data type associated with a deletion operation.
[0051] Optionally, the second type of incremental business data may refer to data with the fourth digit identified as “D”, and the second type of incremental business data may represent “delete the data at the primary key 001 in the 010 table”.
[0052] Optionally, at each moment, the target processing system may compare the deleted data at that moment with the first difference data set and the second difference data set at the previous moment to find inconsistent records and form the second difference data set at that moment.
[0053] Step S304 : determining whether the data in the second business system is abnormal based on the first difference data set and the second difference data set, and obtaining target information.
[0054] Optionally, the target processing system may determine whether there is an anomaly in the data in the second business system based on the first difference data sets at multiple moments and the second difference data sets at multiple moments, and obtain target information.
[0055] For example, based on the first difference data set at each of two adjacent moments and the second difference data set at each of two adjacent moments, it is determined whether there is an anomaly in the data in the second business system to obtain target information.
[0056] For example, for each incremental business data, the "table number + logical primary key" of the data is used as the primary key of the data. When determining the target information, the target processing system can follow the following principles: (1) For data with consistent primary keys, under the add / modify logic, if the data content is consistent (excluding operators), the data of the new and old systems can be considered consistent, and the record of the data in the cache is deleted. (2) If the inconsistent data is still updated in a certain system in the next period (that is, the next moment), it is considered that the data is still changing and the current inconsistency is acceptable. (3) If there is no update in the new and old systems for a certain primary key data, and the content of the last pushed data is inconsistent, the data is considered abnormal (the data cannot be delayed for a long time). (4) If a system modifies the same primary key data twice in a row, and the other system does not have any operation records for both times, the data is considered abnormal. (5) For data deleted by a system, the data recorded in the system itself is first deleted. If there is no data to be deleted (it may have been deleted before because the data is consistent), it is recorded in the deletion backup table of the system. If there is also a deletion record in the other system, the logic is considered correct. Otherwise, the data is considered abnormal.
[0057] Based on the scheme defined in the above steps S301 to S304, it can be known that in an embodiment of the present invention, a method of judging whether the new system has data anomalies based on the incremental business data of the new and old systems at multiple times is adopted. By respectively obtaining the incremental business data of the first business system and the second business system at multiple times, a first business data set of the first business system and a second business data set of the second business system are obtained. Then, the incremental business data of the first type in the first business data set and the second business data set are compared to obtain the first difference data set at each moment. Then, the incremental business data of the second type in the first business data set and the second business data set are compared to obtain the second difference data set at each moment. Thus, based on the first difference data set and the second difference data set, it is judged whether the data in the second business system has anomalies and the target information is obtained. Among them, the second business system is the system to be used to replace the first business system for processing business. The first type refers to the data type associated with the characterization of the new operation or the modification operation, and the second type refers to the data type associated with the characterization of the deletion operation.
[0058] It is easy to notice that in the above process, by comparing the incremental business data belonging to the first type in the first business data set and the second business data set respectively, and comparing the incremental business data belonging to the second type respectively, and performing data anomaly judgment based on the first difference data set and the second difference data set obtained by the comparison, on the one hand, it is possible to determine whether there is an anomaly in the new system data by combining the difference information of the incremental business data of the new and old systems at multiple times, and to consider the data differences between the new and old systems caused by the system processing time, thereby avoiding focusing only on the static consistency of the data, thereby improving the accuracy of judgment. On the other hand, it is possible to distinguish and compare the incremental business data in the new and old systems based on the different data types, thereby further improving the accuracy of judgment.
[0059] It can be seen from this that the solution provided by this application achieves the purpose of judging whether there are data anomalies in the new system based on the incremental business data of the new and old systems at multiple times, thereby achieving the technical effect of improving the accuracy of judging whether there are data anomalies in the new system, and further solving the technical problem of low judgment accuracy when comparing the business data of the new and old business systems in related technologies to judge whether there are data anomalies in the new system.
[0060] In an optional embodiment, the incremental business data belonging to the first type in the first business data set and the second business data set are compared to obtain the first difference data set at each moment, including: determining the incremental business data belonging to the first type in the first business data set as the first business data, and determining the incremental business data belonging to the first type in the second business data set as the second business data; for the first moment among multiple moments, determining the first difference data and the second difference data at the moment based on the first business data at the moment and the second business data at the moment, wherein the first difference data belongs to the first business data and the second difference data belongs to the second business data; for the Nth moment among the multiple moments, determining the first difference data and the second difference data at the moment based on the first business data and the second business data at the moment, and the first difference data and the second difference data at the N-1th moment, wherein N is greater than 1; determining the first difference data set at each moment based on the first difference data and the second difference data at each moment.
[0061] Optionally, for a first moment among the multiple moments, first difference data and second difference data at the moment are determined based on the first service data at the moment and the second service data at the moment.
[0062] For example, the first business data at the first moment is stored in the fourth storage area [Redis:Temp], and then [Redis:Temp] is compared with the second business data at the first moment. For each incremental business data, the "table number bit + logical primary key" of the data is used as the primary key of the data. If the data content (excluding operators) of the data with the same primary key between [Redis:Temp] and the second business data is the same, the two data are considered to be the same. Conversely, if the primary key or data content between the two data is different, the two data are considered to be different. If the two data are determined to be the same, the two data are deleted, and the remaining (i.e., not deleted) data in [Redis:Temp] is determined to be the first difference data, and the remaining (i.e., not deleted) data in the second business data at the first moment is determined to be the second difference data.
[0063] Optionally, the first difference data at that moment may be stored in the fifth storage area [Redis:[T1]Back_Up_Ori] corresponding to that moment, and the second difference data at that moment may be stored in the sixth storage area [Redis:[T1]Back_Up_New] corresponding to that moment.
[0064] Optionally, for the incremental business data at each time point subsequent to the first moment, the target processing system may repeat the operation on the data at the first moment, but this time also taking into account the difference data recorded at the previous moment. For example, for the Nth moment among multiple moments, the data to be compared at that moment is determined based on the first business data at that moment and the first difference data at the N-1th moment. The first difference data and second difference data at that moment are determined based on the data to be compared at that moment, the second business data at that moment, and the second difference data at the N-1th moment, where N is greater than 1.
[0065] For example, the target processing system can check whether there is an update in the first business data (or second business data) at the current time point that matches the second difference data (or first difference data) recorded at the previous time point. For example, if a record was marked as first difference data at the previous time point, then at the current time point, if the new system has not yet recorded this update, the record will continue to be marked as first difference data and may trigger an exception warning. Conversely, if the new system makes up this update at the current time point, then this record will be considered a resolved difference and will not be recorded as the first difference data at the current moment.
[0066] Optionally, for each moment, the first difference data at that moment and the second difference data at that moment are combined into a first difference data set at that moment, and the first difference data set may be an empty set.
[0067] It should be noted that, through the above method, the first difference data set at a certain moment is determined based on the current incremental business data at a certain moment and the difference data at the previous moment, thereby taking into account the data differences caused by different system processing time or business logic updates, and can effectively distinguish between normal differences and actual abnormal data caused by operation delays, data processing logic differences or other time-sensitive factors, thereby improving the accuracy of judging whether there are data anomalies in the new system.
[0068] In an optional embodiment, based on the first business data and the second business data at the moment, and the first difference data and the second difference data at the N-1th moment, the first difference data and the second difference data at the moment are determined, including: determining the primary key of the first business data at the moment as the first target primary key, deleting the data corresponding to the first target primary key in the first difference data at the N-1th moment to obtain the updated first difference data, and determining the first business data at the moment and the updated first difference data as the data to be compared at the moment; deleting the identical data between the data to be compared at the moment and the second difference data at the N-1th moment to obtain the updated data to be compared and the updated second difference data; deleting the identical data between the updated data to be compared and the second business data at the moment to obtain the updated data to be compared again and the updated second business data; determining the updated data to be compared again as the first difference data at the moment, and determining the updated second business data and the updated second difference data as the second difference data at the moment.
[0069] For example, for the Nth moment among multiple moments, the first difference data of the previous moment is queried from [Redis:[T(N-1)]Back_Up_Ori], and the first difference data of the previous moment is stored in [Redis:Temp]. [T(N-1)] is the time point for recording the first difference data (or second difference data) of the N-1th moment. Optionally, the target processing system can record the primary key of these first difference data as "inconsistent data in the new / old system" to identify whether an error needs to be reported later. Afterwards, the first business data of the old system at moment N is parsed, and [Redis:Temp] is added or modified based on the first target primary key. For example, for each first target primary key, if data corresponding to that first target primary key exists in [Redis:Temp], the data corresponding to that first target primary key is deleted from [Redis:Temp] and [Redis:[T(N-1)]Back_Up_Ori], and the record of "existing inconsistent data" for that data primary key in the new system is deleted. The data in [Redis:Temp] after deletion becomes the updated first difference data. The first business data at time N is then stored in [Redis:Temp]. The data in [Redis:Temp] at this point becomes the data to be compared at time N.
[0070] After determining the data to be compared at moment N, compare the data in [Redis:Temp] and [Redis:[T(N-1)]Back_Up_New], that is, compare the data to be compared with the second difference data at moment N-1. If consistent data exists, delete the consistent data from [Redis:Temp] and [Redis:[T(N-1)]Back_Up_New] respectively, and then determine the data in [Redis:Temp] at this time as the updated data to be compared, and determine the data in [Redis:[T(N-1)]Back_Up_New] at this time as the updated second difference data, and delete the record of "existing inconsistent data" of the primary key of the consistent data in the new and old systems.
[0071] Optionally, in this embodiment, for two data belonging to the first type of incremental business data (such as the first business data, second business data, first difference data, and second difference data mentioned above), the same data means that when the primary keys of the two data are the same and the data contents (regardless of operators) are the same, the two data are considered to be the same; otherwise, if the primary keys or data contents between the two data are different, the two data are considered to be different.
[0072] After obtaining the updated data to be compared and the updated second difference data, parse the second business data at moment N, and delete the record of "existing inconsistent data" for the primary key of the data in the old system according to the primary key of the data. Then compare the second business data at moment N with the data in [Redis:Temp]. If there is consistent data, delete the consistent data from [Redis:Temp] and the second business data respectively. Then determine the data in [Redis:Temp] at this time as the updated data to be compared again, and determine the remaining data in the second business data at moment N as the updated second business data, and delete the record of "existing inconsistent data" for the primary key of the consistent data in the old and new systems.
[0073] Optionally, the data currently existing in [Redis:Temp] (ie, the updated data to be compared) is stored in [Redis:[T(N)]Back_Up_Ori], and the updated second business data is stored in [Redis:[T(N)]Back_Up_New].
[0074] It should be noted that, through the above method, effective deletion of identical data and effective determination of different data are achieved, thereby improving the accuracy of the determined first difference data set, thereby improving the accuracy of determining whether there are data anomalies in the new system.
[0075] In an optional embodiment, the incremental business data belonging to the second type in the first business data set and the second business data set are compared to obtain a second difference data set at each moment, including: determining the incremental business data belonging to the second type in the first business data set as the third business data, and determining the incremental business data belonging to the second type in the second business data set as the fourth business data; for the first moment among multiple moments, determining the third difference data and the fourth difference data at the moment based on the third business data at the moment and the fourth business data at the moment, wherein the first difference data belongs to the third business data and the second difference data belongs to the fourth business data; for the Nth moment among the multiple moments, determining the third difference data and the fourth difference data at the moment based on the third business data and the fourth business data at the moment, and the first difference data, second difference data, third difference data and fourth difference data at the N-1th moment; and determining the second difference data set at each moment based on the third difference data and the fourth difference data at each moment.
[0076] Optionally, for a first moment among the multiple moments, the third difference data and the fourth difference data at the moment are determined based on the third service data at the moment and the fourth service data at the moment.
[0077] For example, at the first moment, the third business data is directly stored in the seventh storage area [Redis:[T1]Delete_Back_Up_Ori] corresponding to the moment. Afterwards, it is determined whether [Redis:[T1]Delete_Back_Up_Ori] contains the same data as the fourth business data at the first moment. If so, the data is deleted from [Redis:[T1]Delete_Back_Up_Ori] and the fourth business data, thereby determining the remaining data in [Redis:[T1]Delete_Back_Up_Ori] as the third difference data at the moment, and the remaining data in the fourth business data at the first moment as the fourth difference data at the moment, and storing the fourth difference data in the eighth storage area [Redis:[T1]Delete_Back_Up_New] corresponding to the moment.
[0078] Optionally, in this embodiment, for two data belonging to the second type of incremental business data (such as the third business data, fourth business data, third difference data, and fourth difference data mentioned above), the same data means that when the primary keys of the two data are the same, the two data are considered to be the same; otherwise, if the primary keys between the two data are different, the two data are considered to be different.
[0079] Optionally, for the incremental business data at each time point subsequent to the first moment, the target processing system can repeat the operation on the data at the first moment, but this time it will also take into account the difference data recorded at the previous moment (first difference data, second difference data, third difference data and fourth difference data).
[0080] Optionally, for each moment, the third difference data at that moment and the fourth difference data at that moment are combined into a second difference data set at that moment, and the second difference data set may be an empty set.
[0081] It should be noted that, through the above method, it is possible to determine the second difference data set at a certain moment based on the current incremental business data at a certain moment and the difference data at the previous moment, thereby taking into account the data differences caused by different system processing times or business logic updates, and can effectively distinguish between normal differences and actual abnormal data caused by operation delays, data processing logic differences or other time-sensitive factors, thereby improving the accuracy of judging whether there are data anomalies in the new system.
[0082] In an optional embodiment, the third difference data and the fourth difference data at the moment are determined based on the third business data and the fourth business data at the moment, and the first difference data, the second difference data, the third difference data and the fourth difference data at the N-1th moment, including: determining the primary key of the third business data at the moment as the second target primary key, and when there is no data corresponding to the second target primary key in the first difference data at the N-1th moment and the fourth difference data at the N-1th moment, determining the third business data as the third difference data at the moment; determining the primary key of the fourth business data at the moment as the third target primary key, and when there is no data corresponding to the third target primary key in the second difference data at the N-1th moment and the third difference data at the N-1th moment, determining the fourth business data as the fourth difference data at the moment.
[0083] For example, for the Nth moment among multiple moments, the primary key of the third business data at the moment is determined as the second target primary key, and then based on the second target primary key, first determine whether the data corresponding to the second target primary key is stored in [Redis:[T(N-1)]Back_Up_Ori] (that is, the first difference data at the N-1th moment). If so, delete the data corresponding to the second target primary key and the third business data corresponding to the second target primary key. If not, determine whether the data corresponding to the second target primary key is stored in [Redis:[T(N-1)]Delete_Back_Up_New] (that is, the fourth difference data at the N-1th moment). If so, delete the data and the third business data corresponding to the second target primary key. If not, determine the third business data as the third difference data at the Nth moment and store it in [Redis:[T(N)]Delete_Back_Up_Ori].
[0084] Optionally, the primary key of the fourth business data at the Nth moment is determined as the third target primary key, and then it is determined whether data corresponding to the third target primary key is stored in [Redis:[T(N-1)]Back_Up_New] (i.e., the second difference data at the N-1th moment). If so, the data corresponding to the third target primary key and the fourth business data corresponding to the third target primary key are deleted. If not, it is determined whether data corresponding to the third target primary key is stored in [Redis:[T(N-1)]Delete_Back_Up_Ori] (i.e., the third difference data at the N-1th moment) and [Redis:[T(N)]Delete_Back_Up_Ori] (i.e., the third difference data at the Nth moment). If so, the data and the fourth business data corresponding to the third target primary key are deleted. If not, the fourth business data is determined as the fourth difference data at the Nth moment and stored in [Redis:[T(N)]Delete_Back_Up_New].
[0085] It should be noted that, through the above method, effective deletion of identical data and effective determination of different data are achieved, thereby improving the accuracy of the determined second difference data set, thereby improving the accuracy of determining whether there are data anomalies in the new system.
[0086] In an optional embodiment, based on the first difference data set and the second difference data set, it is determined whether there is an anomaly in the data in the second business system to obtain target information, including: for the Mth moment among multiple moments, recording the first primary key of the moment to the first storage area, wherein the first primary key includes: a primary key for which the first difference data exists but the second difference data does not exist at the moment, and a primary key for which the third difference data exists but the fourth difference data does not exist at the moment, and M is a positive integer; for the M+1th moment among the multiple moments, if the first difference data or the third difference data of the first primary key of the Mth moment in the first storage area is not deleted, and the second business data and the fourth business data identical to the first primary key do not exist at the M+1th moment, then it is determined that the target information indicates that there is an anomaly in the data in the second business system.
[0087] In an optional embodiment, [Back_Up_Ori / New] is used to store temporarily inconsistent data identified during a certain period (i.e., a certain moment) of data push, that is, to store the first difference data and the second difference data; temp is used to temporarily store the inconsistent data of the previous period (i.e., the previous moment) of the old system and the current period data (modification operations on the same primary key data will be directly modified to retain the final situation); [Delete_Back_Up_Ori / New] is used to store data identified during a certain period of data push that cannot be logically deleted within the system and has no corresponding deletion records across systems, that is, to store the third difference data and the fourth difference data.
[0088] In an optional embodiment, the operation record of whether data has been modified is stored in the following collection: Inconsistent1 (i.e., the first storage area) is used to record data from the previous period that only existed in the old system. If there is no deletion record of this data in the old system in the next period, and no new record of this data is added in the new system, this data is counted as Error, which means that the target information indicates that there is an abnormality in the data in the second business system. Records are used to record incremental operations on data (i.e., addition, modification, and deletion). If there are no new records in a period, it can be understood that no incremental operations (i.e., addition, modification, and deletion) were performed on the data in that period.
[0089] For example, for the Mth moment among multiple moments, the primary key for which corresponding data exists in [Redis:[T(M)]Back_Up_Ori] and no corresponding data exists in [Redis:[T(M)]Back_Up_New] is determined as the first primary key; and the primary key for which corresponding data exists in [Redis:[T(M)]Delete_Back_Up_Ori] and no corresponding data exists in [Redis:[T(M)]Delete_Back_Up_New] is determined as the first primary key.
[0090] Optionally, if at the M+1th moment, the first difference data or the third difference data for the first primary key at the Mth moment in the first storage area has not been deleted, and the second business data and the fourth business data identical to the first primary key do not exist at the M+1th moment, that is, if there is no deletion record of the data for the first primary key in the old system in the next period, and the new system does not have a new record of the data for the first primary key, then the data is counted as an Error, and it is determined that the target information indicates that there is an abnormality in the data in the second business system. It should be noted that, at any moment, the first difference data and the third difference data corresponding to the same first primary key will not exist simultaneously in the first storage area.
[0091] In an optional embodiment, when it is not determined that the data in the second business system has an anomaly, it may be determined that the target information indicates that the data in the second business system has no anomaly.
[0092] It should be noted that, through the above method, data comparison across time points is achieved, thereby improving the accuracy of determining target information.
[0093] In an optional embodiment, based on the first difference data set and the second difference data set, it is determined whether there is an anomaly in the data in the second business system to obtain target information, including: for the Mth moment among multiple moments, the second primary key of the moment is recorded to the second storage area, wherein the second primary key includes: a primary key for which the second difference data exists but the first difference data does not exist at the moment, and a primary key for which the fourth difference data exists but the third difference data does not exist at the moment, and M is a positive integer; for the M+1th moment among the multiple moments, if the second difference data or the fourth difference data of the second primary key of the Mth moment in the second storage area is not deleted, and the first business data and the third business data identical to the second primary key do not exist at the M+1th moment, then it is determined that the target information indicates that there is an anomaly in the data in the second business system.
[0094] In an optional embodiment, the operation record of whether the data is modified is stored in the collection as follows: Inconsistent2 (i.e., the second storage area) is used to record the data of the previous period that only exists in the new system. If there is no deletion record of the data by the new system in the next period, and the old system has no new record of the data, the data will be counted as Error, that is, it is determined that the target information represents that there is an abnormality in the data in the second business system.
[0095] For example, for the Mth moment among multiple moments, the primary key for which corresponding data exists in [Redis:[T(M)]Back_Up_New] and no corresponding data exists in [Redis:[T(M)]Back_Up_Ori] is determined as the second primary key; and the primary key for which corresponding data exists in [Redis:[T(M)]Delete_Back_Up_New] and no corresponding data exists in [Redis:[T(M)]Delete_Back_Up_Ori] is determined as the second primary key.
[0096] Optionally, if at the M+1th moment, the second difference data or the fourth difference data for the second primary key at the Mth moment in the first storage area has not been deleted, and the first business data and the third business data identical to the second primary key do not exist at the M+1th moment, that is, if there is no record of deletion of the data for the second primary key by the new system in the next period, and the old system does not have a new record of adding the data for the second primary key, then the data is counted as an Error, and it is determined that the target information indicates that there is an abnormality in the data in the second business system. It should be noted that, at any given moment, the second difference data and the fourth difference data corresponding to the same second primary key will not exist simultaneously in the first storage area.
[0097] In an optional embodiment, when it is not determined that the data in the second business system has an anomaly, it may be determined that the target information indicates that the data in the second business system has no anomaly.
[0098] It should be noted that, through the above method, data comparison across time points is achieved, thereby improving the accuracy of determining target information.
[0099] In an optional embodiment, based on the first difference data set and the second difference data set, it is determined whether there is an anomaly in the data in the second business system to obtain target information, including: for the Mth moment among multiple moments, the third primary key of the moment is recorded to the third storage area, wherein the third primary key includes: the primary key of one of the first difference data and the third difference data exists at the moment, and the primary key of one of the second difference data and the fourth difference data exists, and M is a positive integer; if there is no incremental business data corresponding to the third primary key in the third storage area at the M+1th moment, it is determined that the target information indicates that there is an anomaly in the data in the second business system.
[0100] In an optional embodiment, the operation record of whether the data is modified is stored in the collection as follows: Inconsistent3 (i.e., the third storage area) is used to record the data of the previous period that exists in both the old and new systems but is inconsistent. If there is no new record of the data in the next period in both systems, it is considered that the data has reached a "local final" state within the current time slice, that is, the user has stopped continuously modifying the data. If the data content is still inconsistent at this time, the data will be counted as Error, that is, it is determined that the target information represents that there is an abnormality in the data in the second business system.
[0101] For example, for the Mth moment among multiple moments, the primary key for which corresponding data exists in one of [Redis:[T(M)]Back_Up_Ori], [Redis:[T(M)]Delete_Back_Up_Ori] and corresponding data exists in one of [Redis:[T(M)]Back_Up_New], [Redis:[T(M)]Delete_Back_Up_New] is determined as the third primary key.
[0102] Optionally, if at the M+1th moment, there is no incremental business data corresponding to the third primary key in the third storage area (i.e., the first business data, the second business data, the third business data, and the fourth business data), that is, both systems have no new records of the data, then it is determined that the target information indicates that there is an abnormality in the data in the second business system.
[0103] In an optional embodiment, when no abnormality is determined in the data in the second business system under any of the above situations (i.e., situations corresponding to the first storage area, the second storage area, and the third storage area), it can be determined that the target information indicates that there is no abnormality in the data in the second business system.
[0104] It should be noted that, through the above method, data comparison across time points is achieved, thereby improving the accuracy of determining target information.
[0105] In an optional embodiment, by Figure 4 The example in illustrates the logic of determining target information in this application. Figure 4 is a schematic diagram of an optional comparison of incremental business data according to an embodiment of the present invention. Figure 4 Where L[i] represents the actual i-th modification (addition / modification / deletion) of the data, T[i] represents the i-th push of the system (equivalent to the i-th moment), I represents the addition of pushed data, C represents the modification of pushed data, and D represents the deletion of pushed data. [Back_Up_Ori / New] is used to store temporarily inconsistent data identified during a certain period of data push, temp is used to temporarily store inconsistent data from the previous period of the old system and the current period's data (modification operations on data with the same primary key will be directly modified to retain the final situation), and [Delete_Back_Up_Ori / New] is used to store data identified during a certain period of data push that cannot be logically deleted within the system and has no corresponding deletion records across systems. As shown in 4, this example only takes two or three consecutive modifications of data as an example, and the actual situation is not limited to this:
[0106] Case 0 [Most Basic]: Two rows of data from the same period are parsed and matched. After the old file data is stored in the temp file, the new file is compared based on the primary key and the original data is deleted. After the comparison is complete, there is no record of this data in the temp file, so it is not shown in the diagram.
[0107] Case 1 [Normal]: The user adds a new data record L1 and then modifies it to record L2. After that, the data temporarily stops changing. During the push of the T1 batch data file (i.e., incremental business data), there is no previous record of this data. The old system data has completed the recording of the second data change. The data-added state L2I is pushed to the downstream system. The target processing system parses the file and records the data in temp. When the new system pushes, there is no previous record of this data. The data has not completed the recording of the change from the first state to the second state. Therefore, the data-added state L1I is pushed to the downstream. Redis parses the file and finds that the two phases of data are inconsistent. The record of the old system in temp (i.e., the record used to record the new operation on the data, which is equivalent to the first difference data) is stored in [T1]Back_Up_Ori. At the same time, the record of the new system (i.e., the record used to record the new operation on the data, which is equivalent to the second difference data) is stored in [T1]Back_Up_New for comparison in the next phase. During the push of the T2 batch data file, the target processing system first reads the record of the data in [T1]Back_Up_Ori and stores it in temp. Since the push process of the previous period has completed the push of the data L2, the old system in this period has no further push for the data, and temp is not updated. Compared with the previous record [T1]Back_Up_New of the new system, it cannot be eliminated. In this stage, the new system recognizes that the data has changed and pushes L2C. At this time, the data pushed by the new system is consistent with the data in temp, and the record of the data in temp and the data pushed by the new system is deleted.
[0108] Case 2 [Normal]: Similar to Case 1, the user adds data record L1, then modifies it to record L2, and the data temporarily stops changing. During the T1 batch data file push process, there was no previous record of this data, and the old system data had not yet completed recording the second data change. The data with the newly added status L1I is pushed to the downstream system. The target processing system parses the file and records the data in temp. When the new system pushes, there was no previous record of this data, and the second data change has been recorded. Therefore, the data with the newly added status L2I is pushed to the downstream. The target processing system parses the file and finds that the two data periods are inconsistent. The record of the old system in temp is stored in [T1]Back_Up_Ori, and the data of the new system is stored in [T1]Back_Up_New for comparison in the next period. During the push of the T2 batch data file, Redis first reads the record for the data in [T1]Back_Up_Ori and stores it in temp. It then identifies the update L2C for the data pushed in this period's data file, updates the data in temp, compares it with the previous period's record [T1]Back_Up_New in the new system, and deletes the corresponding records. The new system does not push any data changes for this period.
[0109] Case 3 [Normal situation]: The user adds new data L1, modifies it to L2, and then modifies it to L3. When the T1 batch is pushed, the old system recognizes the second modification status of the data and pushes L2I. The new system only recognizes the first modification and pushes L1I. The target processing system handles the same situation as Case 1, and T1 ends. When the T2 batch is pushed, the target processing system first reads the record of the data in [T1]Back_Up_Ori and stores it in temp. The old system data file pushes the update L3C of the data, updates the data stored in temp, and then compares it with the record in [T1]Back_Up_New. If it is inconsistent, it is temporarily stored for subsequent comparison processes. The new system data file pushes the update L2C of the data. If it is inconsistent with the data in temp, it is recorded in [T2]Back_Up_Ori and [T2]Back_Up_New respectively. When the T3 batch is pushed, the new system data is updated to the L3 status. During the comparison process of the target processing system, the data of the two systems are consistent, and the record of the data is deleted.
[0110] Case 4 [Normal]: The user adds data to L1, modifies it to L2, and then modifies it to L3. After the T1 batch is pushed, Redis processes the data in the same way as in Case 1. After the T2 batch is pushed, the data in both systems is synchronized to the L3 state, and the data can be eliminated.
[0111] Case 5 [Normal]: The user adds data in L1, modifies it to L2, and then modifies it to L3. After batch 1 is pushed, Redis processes the data in the same way as in Case 1. Batch T2 is pushed, and the new system recognizes the data as being in L3. The old system's data remains in L2. The temp file is updated, recording [T2]Back_Up_Ori (data in L2 at T1) and [T2]Back_Up_New (data in L3 at T2). Batch T3 is pushed, and the old system recognizes the data as being in L3. Both systems synchronize their data to L3, eliminating the issue.
[0112] Case 6 [Normal situation]: The user adds data L1 and then deletes the data. The old system may have completed the addition and deletion of the data before the T1 batch push, or between the T1 batch and the T2 batch push, so there is no push record for the data. The new system T1 batch only identifies the new record of the data and pushes L1I. During the comparison process, Redis does not find data consistent with the primary key and records [T1]Back_Up_New. When the T2 batch is pushed, the new system pushes the deletion of the data, queries [T1]Back_Up_New for records of the data, executes the deletion logic in the system, deletes the record of the data in [T1]Back_Up_New, and simultaneously attempts to delete the deletion record (empty) of the data in [T1]Delete_Back_Up_Ori.
[0113] Case 7 [Normal]: Similar to Case 6, the user adds data L1 and then deletes it. The new system completes the record operations from adding to deleting this data within a single push batch cycle, while the old system spans two batches. After the T1 batch is pushed, Redis records [T1]Back_Up_Ori. When the T2 batch is pushed, [T1]Back_Up_Ori is stored in temp, the system's deletion logic is executed, and the records containing this data in temp are queried. The system's deletion logic is executed, deleting the records in temp for this data. Simultaneously, an attempt is made to delete the deletion record (empty) for this data in [T1]Delete_Back_Up_New.
[0114] Case 8 [Normal]: A user adds data L1 and then deletes it. During the T1 batch push, the new system recognizes the new data addition L1 and pushes it, with Redis recording [T1]Back_Up_New. During the T2 batch push, the old system recognizes the new data addition L1 and pushes it. After storing it in temp, it compares it with the [T1]Back_Up_New data and deletes it. The new system pushes the delete operation on L1 and tries the system's deletion logic. Since there is no [T1]Back_Up_New record for this data, the deletion record is organized into [T2]Delete_Back_Up_New for subsequent comparison. During the T3 batch push, the old system pushes the delete operation for the data and tries the deletion logic within the system. At this time, there is no record [T2]Back_Up_Ori about the data in Redis. When trying to compare between systems, a record [T2]Delete_Back_Up_New can be found. The two records are compared only for consistency in the primary key (deleting data only compares the logical primary key because it is possible to delete data before and after the modification of the same record, which is logically correct). The corresponding Redis records are deleted respectively.
[0115] Case 9 [Normal]: The user adds data L1 and then deletes it. During the T1 batch push, the old system recognizes the data addition operation L1 and pushes it, with Redis recording [T1]Back_Up_Ori. During the T2 batch push, the data in [T1]Back_Up_Ori is first stored in temp. At this point, the deletion operation on the data is recognized, and a query is made in temp to find records of the data. The system's deletion logic is executed, deleting the records in temp and simultaneously attempting to delete the empty deletion record in [T1]Delete_Back_Up_New. The new system then pushes the data addition operation L1. Since there are no related records in Redis, the data operation is stored in [T2]Back_Up_New. During the T3 batch push, the new system pushes the data deletion operation, executing the system's deletion operation.
[0116] Case 10 [Abnormal Situation]: The old system pushed data for a certain primary key, and no records were deleted during the continuous push interval. The new system did not push the data for this primary key. During the T1 batch push, the old system pushed the data, which was recorded as [T1]Back_Up_Ori, and the primary key was recorded as Inconsistent1. During the T2 batch push, neither the old nor the new system had any processing records for this data. That is, the data corresponding to the primary key in Inconsistent1 was not deleted at T2, and there was no corresponding update in the new system. Therefore, this data was recorded as Error, and the error was subsequently reported uniformly through the early warning device.
[0117] Case 11 [Abnormal Situation]: The new system pushed data for a specific primary key, and no records were deleted during the continuous push interval. The old system did not push the data for this primary key. During the T1 batch push, the new system pushed the data, which was recorded as [T1]Back_Up_New, and the primary key was recorded as Inconsistent2. During the T2 batch push, neither the new nor the old system had any processing records for this data. That is, the data corresponding to the primary key in Inconsistent2 was not deleted at T2, and there was no corresponding update in the old system. This data was recorded as Error, and the error was subsequently reported uniformly through the early warning system.
[0118] Case 12 [Abnormal Situation]: The new and old systems pushed different data, and neither system updated the primary key data within the consecutive push interval. During the T1 batch push, the inconsistent data pushed by the new and old systems was recorded as [T1]Back_Up_Ori and [T1]Back_Up_New, respectively, and the primary key was recorded as Inconsistent3. During the T2 batch push, no new records were added. That is, the primary key in Inconsistent3 had no corresponding update in the new and old systems at T2. This was recorded as Error, and the error was subsequently reported uniformly through the early warning system.
[0119] Case 13 [Abnormal Situation]: The final data pushed between the old and new systems during a local time interval is inconsistent. During the T1 batch push, the old system pushed data L1, which was recorded in [T1]Back_Up_Ori, and the primary key was recorded in Inconsistent1. During the T2 batch push, the new system pushed data L2, which was recorded in [T2]Back_Up_New. Due to the inconsistency between the current data between the old and new systems, the old system data was recorded in [T2]Back_Up_Ori, and the primary key was transferred from Inconsistent1 to Inconsistent3. During the T3 batch push, no new records were added, and the data was recorded as Error. The error was subsequently reported uniformly through the early warning system.
[0120] Case 14 [Exception]: Missing deletion in the new system. During the T1 batch push, both the old and new systems pushed data L1. The data was deleted normally, and Redis had no related records. During the T2 batch push, the old system pushed a deletion record for this data. Redis previously had no record of this data, so this was recorded as [T2]Delete_Back_Up_Ori, and this primary key was recorded as Inconsistent1. During the T3 batch push, the other system had no deletion record for this data, so this was recorded as Error, and subsequently reported to the early warning system.
[0121] Case 15 [Exception]: The new system has additional deletion records. During the T1 batch push, both the old and new systems pushed data L1. The data was deleted normally, and Redis had no related records. During the T2 batch push, the new system pushed a deletion record for this data. Redis previously had no record of this data, so this was recorded as [T2]Delete_Back_Up_New, and this primary key was recorded as Inconsistent2. During the T3 batch push, the other system had no deletion record for this data, so this was recorded as Error, and subsequently reported to the early warning system.
[0122] In an optional embodiment, by Figure 5 An optional application process of this embodiment is described. Figure 5 This is a flow chart of an optional data processing method according to an embodiment of the present invention. Figure 2 .like Figure 5 As shown, the target processing system includes a file comparison tool module and an abnormal warning module. The file comparison tool module is divided into a full file comparison submodule and an incremental file comparison submodule, which compare the corresponding file contents respectively. The target processing system can obtain the first business system (i.e. Figure 5 The old system in the second business system (ie Figure 5 The new system in the system) has incremental business data at multiple times (i.e. Figure 5 ), and obtain the full amount of business data of the first business system and the second business system at multiple times (i.e. Figure 5 For each moment, the full file at that moment is compared using Myers differential comparison, parsing the data line by line and comparing the paths that have been modified to achieve a consistent result with the least number of steps. For the incremental files at that moment, the headers are strictly compared for consistency, and logically inconsistent data is temporarily stored in the cache for subsequent comparison to delete logical consistency or to warn of problems. The overall comparison process is as follows:
[0123] (1) Compare the results of the file generation at the specified time of the old system to see if there are any extra full files or missing full files in the new system. If so, report an error; if not, proceed to the next step.
[0124] (2) Compare the MD5 codes of the files in the two systems. If the MD5 codes are consistent, it means that the data in this period is consistent and the full file comparison is completed. For incremental files, the difference data between the two systems in this period (i.e., the current time) and the difference data between the two systems in the previous period (i.e., the previous time) are used to determine whether an error needs to be reported.
[0125] (3) If the MD5 codes are inconsistent, the full file is compared using the Myers differential algorithm, and the comparison results generate a differential file. The incremental file can first be strictly compared to see if the headers are consistent. If they are inconsistent, an error is reported. If they are consistent, the data is then parsed line by line to determine and eliminate the difference data (i.e. Figure 5 The inconsistent data of this period will be temporarily stored, and the abnormal situation of the previous period will be warned (that is, whether to warn based on the target information).
[0126] Optionally, if the amount of data in the differential file corresponding to the full file of the same period exceeds a preset amount, an early warning may be issued. Otherwise, if it is determined to be a normal phenomenon caused by time differences such as data processing, no early warning may be issued. In other words, in this embodiment, whether to issue an early warning depends mainly on the comparison of the incremental files.
[0127] In an optional embodiment, the full set of files can be compared using the Myers differential algorithm. The Myers differential algorithm can calculate and record the minimum number of deletion / addition steps from file 1 to file 2. Figure 6 FIG. 1 is a schematic diagram of an optional method for comparing data using a differential algorithm according to an embodiment of the present invention. Figure 6 , which are the comparison process and results of two files named "ABCDEFGH" and "ABCDEHG". Myers' differential algorithm can record the shortest editing path from file 1 to file 2 by drawing an editing graph from the starting point to the end point. It includes the following steps: (1) Take the hash value of each line of data in the file as a data element, such as Figure 6 Any single capital letter in represents a data element. (2) The two file pointers point to the starting position of their respective files, and the shortest path search is performed from the upper left to the lower right. Movement to the right or downward represents a change in data and is counted as a step. Diagonal movement indicates that the two data are consistent and is not counted as a step. Figure 6 (3) Inversely infer the shortest path generation and convert it into deletion / addition modification operations, such as Figure 6 In (ABCD+EF-HG+H), a difference file is written for the full file comparison. For example, if the original data queue length is N and the target data queue length is M, the computational space is size = 2*(M+N+1). The Diff[] array stores the optimal position for each step and each offset k. In the outer loop, the number of steps is d <= M+N, d = i. In the inner loop, the value of k is in the range [-i, i], with a step size of 2. If the current position data of the two queues is the same, move to the diagonal: i++, j++. Upon reaching the target position, all predecessor nodes stored in the Diff[] array are the optimal path solutions.
[0128] It can be seen from this that the solution provided by this application achieves the purpose of judging whether there are data anomalies in the new system based on the incremental business data of the new and old systems at multiple times, thereby achieving the technical effect of improving the accuracy of judging whether there are data anomalies in the new system, and further solving the technical problem of low judgment accuracy when comparing the business data of the new and old business systems in related technologies to judge whether there are data anomalies in the new system.
[0129] Example 2
[0130] According to an embodiment of the present invention, an embodiment of a data processing device is provided, wherein: Figure 7 is a schematic diagram of an optional data processing device according to an embodiment of the present invention, such as Figure 7 As shown, the device includes:
[0131] An acquisition module 701 is configured to acquire incremental business data of a first business system and a second business system at multiple time points, respectively, to obtain a first business data set of the first business system and a second business data set of the second business system, wherein the second business system is the system that is to replace the first business system in processing business;
[0132] A first comparison module 702 is configured to compare incremental business data of a first type in the first business data set and the second business data set, respectively, to obtain a first difference data set at each moment, wherein the first type refers to a data type associated with a new operation or a modification operation;
[0133] A second comparison module 703 is configured to compare incremental business data of the second type in the first business data set and the second business data set, respectively, to obtain a second difference data set at each moment, wherein the second type refers to a data type associated with a deletion operation;
[0134] The processing module 704 is configured to determine whether there is an anomaly in the data in the second business system based on the first difference data set and the second difference data set, and obtain target information.
[0135] It should be noted that the above-mentioned acquisition module 701, first comparison module 702, second comparison module 703 and processing module 704 correspond to steps S301 to S304 in the above-mentioned embodiment. The examples and application scenarios implemented by the four modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned embodiment 1.
[0136] Optionally, the first comparison module also includes: a first determination submodule, used to determine the incremental business data of the first type in the first business data set as the first business data, and to determine the incremental business data of the first type in the second business data set as the second business data; a second determination submodule, used to determine, for the first moment among multiple moments, the first difference data and the second difference data at the moment based on the first business data at the moment and the second business data at the moment, wherein the first difference data belongs to the first business data and the second difference data belongs to the second business data; a third determination submodule, used to determine, for the Nth moment among the multiple moments, the first difference data and the second difference data at the moment based on the first business data and the second business data at the moment, and the first difference data and the second difference data at the N-1th moment, wherein N is greater than 1; a fourth determination submodule, used to determine the first difference data set for each moment based on the first difference data and the second difference data at each moment.
[0137] Optionally, the third determination submodule also includes: a first determination unit, used to determine the primary key of the first business data at that moment as the first target primary key, delete the data corresponding to the first target primary key in the first difference data at the N-1th moment, obtain the updated first difference data, and determine the first business data at that moment and the updated first difference data as the data to be compared at that moment; a first deletion unit, used to delete the identical data between the data to be compared at that moment and the second difference data at the N-1th moment, obtain the updated data to be compared and the updated second difference data; a second deletion unit, used to delete the identical data between the updated data to be compared and the second business data at that moment, obtain the updated data to be compared again and the updated second business data; a second determination unit, used to determine the updated data to be compared again as the first difference data at that moment, and determine the updated second business data and the updated second difference data as the second difference data at that moment.
[0138] Optionally, the second comparison module includes: a fifth determination submodule, used to determine the incremental business data belonging to the second type in the first business data set as the third business data, and to determine the incremental business data belonging to the second type in the second business data set as the fourth business data; a sixth determination submodule, used to determine, for the first moment among multiple moments, the third difference data and the fourth difference data at the moment based on the third business data at the moment and the fourth business data at the moment, wherein the first difference data belongs to the third business data and the second difference data belongs to the fourth business data; a seventh determination submodule, used to determine, for the Nth moment among the multiple moments, the third difference data and the fourth difference data at the moment based on the third business data and the fourth business data at the moment, and the first difference data, second difference data, third difference data and fourth difference data at the N-1th moment; an eighth determination submodule, used to determine the second difference data set for each moment based on the third difference data and the fourth difference data at each moment.
[0139] Optionally, the seventh determination submodule also includes: a third determination unit, used to determine the primary key of the third business data at the moment as the second target primary key, and when there is no data corresponding to the second target primary key in the first difference data at the N-1th moment and the fourth difference data at the N-1th moment, determine the third business data as the third difference data at the moment; a fourth determination unit, used to determine the primary key of the fourth business data at the moment as the third target primary key, and when there is no data corresponding to the third target primary key in the second difference data at the N-1th moment and the third difference data at the N-1th moment, determine the fourth business data as the fourth difference data at the moment.
[0140] Optionally, the processing module also includes: a first processing sub-module, for recording the first primary key of the Mth moment among multiple moments to the first storage area, wherein the first primary key includes: a primary key for which the first difference data exists but the second difference data does not exist at the moment, and a primary key for which the third difference data exists but the fourth difference data does not exist at the moment, and M is a positive integer; a ninth determination sub-module, for which, for the M+1th moment among the multiple moments, if the first difference data or the third difference data of the first primary key of the Mth moment in the first storage area is not deleted, and the second business data and the fourth business data identical to the first primary key do not exist at the M+1th moment, then determining that the target information represents that there is an abnormality in the data in the second business system.
[0141] Optionally, the processing module also includes: a second processing sub-module, for recording the second primary key of the Mth moment among multiple moments to the second storage area, wherein the second primary key includes: a primary key for which the second difference data exists but the first difference data does not exist at the moment, and a primary key for which the fourth difference data exists but the third difference data does not exist at the moment, and M is a positive integer; a tenth determination sub-module, for which, for the M+1th moment among the multiple moments, if the second difference data or the fourth difference data of the second primary key of the Mth moment in the second storage area is not deleted, and the first business data and the third business data identical to the second primary key do not exist at the M+1th moment, then determining that the target information represents that there is an abnormality in the data in the second business system.
[0142] Optionally, the processing module also includes: a third processing sub-module, used to record the third primary key of the Mth moment among multiple moments to the third storage area, wherein the third primary key includes: a primary key at which one of the first difference data and the third difference data exists at the moment, and one of the second difference data and the fourth difference data exists, and M is a positive integer; an eleventh determination sub-module, used to determine that the target information represents that there is an abnormality in the data in the second business system if there is no incremental business data corresponding to the third primary key in the third storage area at the M+1th moment.
[0143] Example 3
[0144] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is further provided, in which a computer program is stored. The computer program is configured to execute the above-mentioned data processing method when running.
[0145] Example 4
[0146] According to another aspect of an embodiment of the present invention, an electronic device is provided, wherein: Figure 8 is a schematic diagram of an optional electronic device according to an embodiment of the present invention, such as Figure 8 As shown, the electronic device includes one or more processors; a memory for storing one or more programs, which, when the one or more programs are executed by the one or more processors, enables the one or more processors to run the programs, wherein the programs are configured to execute the above-mentioned data processing method when running.
[0147] Example 5
[0148] According to another aspect of an embodiment of the present invention, a computer program product is provided, including a computer program / instruction, which implements the above-mentioned data processing method when the computer program / instruction is executed by a processor.
[0149] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.
[0150] In the above embodiments of the present invention, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0151] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0152] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected to achieve the purpose of the present embodiment according to actual needs.
[0153] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0154] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk, etc. Various media that can store program codes.
[0155] The above are only preferred embodiments of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A data processing method, characterized in that: The method comprises: Acquire incremental business data of the first business system and the second business system at multiple times, respectively, to obtain a first business data set of the first business system and a second business data set of the second business system, wherein the second business system is a system to replace the first business system in processing business; Comparing incremental business data of a first type in the first business data set and the second business data set, respectively, to obtain first difference data sets at each moment, wherein the first type refers to a data type associated with a new operation or a modification operation; Comparing incremental business data of the second type in the first business data set and the second business data set, respectively, to obtain second difference data sets at each moment, wherein the second type refers to a data type associated with a deletion operation; Based on the first difference data set and the second difference data set, it is determined whether there is an anomaly in the data in the second business system, and target information is obtained.
2. The method according to claim 1, characterized in that Comparing incremental service data of the first type in the first service data set and the second service data set to obtain first difference data sets at each moment, including: Determine the incremental service data of the first type in the first service data set as first service data, and determine the incremental service data of the first type in the second service data set as second service data; For a first moment among the multiple moments, determining first difference data and second difference data at the moment based on the first service data at the moment and the second service data at the moment, wherein the first difference data belongs to the first service data and the second difference data belongs to the second service data; For an Nth moment among the multiple moments, determining first difference data and second difference data at the moment based on the first service data and the second service data at the moment, and the first difference data and the second difference data at the N-1th moment, where N is greater than 1; A first difference data set at each moment is determined based on the first difference data and the second difference data at each moment.
3. The method according to claim 2, characterized in that Determining the first difference data and the second difference data at the moment based on the first service data and the second service data at the moment, and the first difference data and the second difference data at the N-1th moment, includes: Determine the primary key of the first business data at the time as the first target primary key, delete the data corresponding to the first target primary key in the first difference data at the N-1th time, obtain updated first difference data, and determine the first business data at the time and the updated first difference data as the data to be compared at the time; Deleting the identical data between the data to be compared at that moment and the second difference data at the N-1th moment, to obtain updated data to be compared and updated second difference data; Deleting identical data between the updated data to be compared and the second business data at that moment, to obtain updated data to be compared and updated second business data; The updated data to be compared is determined as the first difference data at this moment, and the updated second business data and the updated second difference data are determined as the second difference data at this moment.
4. The method according to claim 3, characterized in that Comparing the incremental service data of the second type in the first service data set and the second service data set to obtain a second difference data set at each moment, including: Determine the incremental service data of the second type in the first service data set as third service data, and determine the incremental service data of the second type in the second service data set as fourth service data; For a first moment among the multiple moments, determining third difference data and fourth difference data at the moment based on the third service data and the fourth service data at the moment, wherein the first difference data belongs to the third service data and the second difference data belongs to the fourth service data; For an Nth moment among the multiple moments, determining third difference data and fourth difference data at the moment based on the third service data and the fourth service data at the moment, and the first difference data, the second difference data, the third difference data, and the fourth difference data at the N-1th moment; A second difference data set at each moment is determined based on the third difference data and the fourth difference data at each moment.
5. The method according to claim 4, characterized in that Determining the third difference data and the fourth difference data at the moment based on the third service data and the fourth service data at the moment, and the first difference data, the second difference data, the third difference data, and the fourth difference data at the N-1th moment, includes: Determine the primary key of the third business data at the time as the second target primary key; if no data corresponding to the second target primary key exists in the first difference data at the N-1th time and the fourth difference data at the N-1th time, determine the third business data as the third difference data at the time; The primary key of the fourth business data at this moment is determined as the third target primary key. When there is no data corresponding to the third target primary key in the second difference data at the N-1th moment and the third difference data at the N-1th moment, the fourth business data is determined to be the fourth difference data at this moment.
6. The method according to claim 4, characterized in that Determining whether data in the second business system is abnormal based on the first difference data set and the second difference data set, and obtaining target information, includes: For an Mth moment among the multiple moments, recording a first primary key of the moment in the first storage area, wherein the first primary key includes: a primary key for which the first difference data exists but the second difference data does not exist at the moment, and a primary key for which the third difference data exists but the fourth difference data does not exist at the moment, and M is a positive integer; For the M+1th moment among the multiple moments, if the first difference data or the third difference data of the first primary key at the Mth moment in the first storage area is not deleted, and there is no second business data and fourth business data that are the same as the first primary key at the M+1th moment, then it is determined that the target information represents that there is an abnormality in the data in the second business system.
7. The method according to claim 4, characterized in that Determining whether data in the second business system is abnormal based on the first difference data set and the second difference data set, and obtaining target information, includes: For an Mth moment among the multiple moments, recording a second primary key of the moment in the second storage area, wherein the second primary key includes: a primary key for which the second difference data exists but the first difference data does not exist at the moment, and a primary key for which the fourth difference data exists but the third difference data does not exist at the moment, and M is a positive integer; For the M+1th moment among the multiple moments, if the second difference data or the fourth difference data of the second primary key at the Mth moment in the second storage area is not deleted, and there is no first business data and third business data identical to the second primary key at the M+1th moment, then it is determined that the target information represents that there is an abnormality in the data in the second business system.
8. The method according to claim 4, characterized in that Determining whether data in the second business system is abnormal based on the first difference data set and the second difference data set, and obtaining target information, includes: For an Mth moment among the multiple moments, recording a third primary key of the moment in the third storage area, wherein the third primary key comprises: a primary key for which one of the first difference data and the third difference data exists at the moment, and one of the second difference data and the fourth difference data exists, and M is a positive integer; If there is no incremental business data corresponding to the third primary key in the third storage area at the M+1th moment, it is determined that the target information indicates that there is an abnormality in the data in the second business system.
9. A data processing device, characterized in that: The device comprises: an acquisition module, configured to respectively acquire incremental business data of the first business system and the second business system at multiple time points, to obtain a first business data set of the first business system and a second business data set of the second business system, wherein the second business system is a system to replace the first business system in processing business; a first comparison module, configured to compare incremental business data of a first type in the first business data set and the second business data set, respectively, to obtain a first difference data set at each moment, wherein the first type refers to a data type associated with a new operation or a modification operation; a second comparison module, configured to compare incremental business data of a second type in the first business data set and the second business data set, respectively, to obtain a second difference data set at each moment, wherein the second type refers to a data type associated with a deletion operation; The processing module is used to determine whether there is an anomaly in the data in the second business system based on the first difference data set and the second difference data set, and obtain target information.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program is configured to execute the data processing method according to any one of claims 1 to 8 when executed.
11. An electronic device, characterized in that: The electronic device includes one or more processors; A memory for storing one or more programs, which, when executed by the one or more processors, enables the one or more processors to run the programs, wherein the programs are configured to execute the data processing method described in any one of claims 1 to 8 when run.
Citation Information
Patent Citations
Incremental data synchronization method and device, computer equipment and storage medium
CN111538779A
System and method for realizing incremental data comparison
CN112231324A
Processing method for incremental data synchronization
CN116975159A
Data comparison method and device, computer equipment and storage medium
CN119415494A