A data migration verification method and system for an information technology application innovation platform
By adopting multi-dimensional data checks and dynamic monitoring methods on the information innovation platform, the data consistency problem in the database migration process in the existing technology is solved, and high reliability and consistency in the data migration process is achieved.
Patent Information
- Application Number
- CN202510060872.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-01-15
AI Technical Summary
The existing database migration technology ignores the relationship between data in different dimensions and dynamic data changes in the data verification process, resulting in the problem of data integrity and consistency after migration cannot be discovered in a timely manner.
A data migration verification method for the information creation platform is proposed. By checking and dynamically monitoring the multi-dimensional characteristics such as labels, basic information, similarity analysis, etc. of the migrated data, using a similarity matching algorithm, dynamic data table check-up and target source data comparison, ensuring high reliability and consistency during the data migration process.
Real-time monitoring and accurate identification of potential data problems during data migration process is achieved, ensuring high reliability and consistency of data migration, and minimizing errors and data loss during data migration process.
Smart Images

Figure CN119474078B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to a method and system for data migration verification in an information and communication technology (ICT) innovation platform. Background Art
[0002] With the development of enterprises, the performance and storage capacity of databases often cannot meet the growing business needs. To adapt to the needs of business expansion, enterprises usually perform database migration, transferring data from the original database to a new database. During the migration process, the original database and the target database may be located in different geographical locations or computer rooms, and the business system usually runs in real time. The data migration needs to be carried out without interrupting the business.
[0003] In the data verification link of existing database migration technologies, data loss is mainly avoided through step-by-step verification and backup of migration path nodes. However, data migration not only involves the transmission of data itself, but also involves verification in multiple dimensions such as data tags, migration tables, data source consistency, and similarity analysis. These traditional methods often ignore the relationships between data in different dimensions, upstream and downstream dependencies, and dynamic data changes during the migration process, resulting in the inability to timely detect data integrity and consistency problems after migration, and these problems often only become apparent after the migration is completed.
[0004] Therefore, how to use a multi-dimensional verification method to monitor and accurately identify potential data problems in real time during the migration process has become an urgent technical problem to be solved. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to propose a method and system for data migration verification in an ICT innovation platform, which can, during the data migration process, verify and dynamically monitor multi-dimensional features such as the tags, basic information, and similarity analysis of the migrated data, and automatically identify and correct data missing or inconsistent problems during the real-time data migration process. This method ensures high reliability and consistency during the data migration process by introducing a similarity matching algorithm, dynamic data table verification, and comparison with target source data.
[0006] To achieve the above purpose, the present invention provides the following technical solutions:
[0007] Based on the above purpose, in the first aspect, the present invention provides a method for data migration verification in an ICT innovation platform, including the following steps:
[0008] Read all source data to be migrated in the ICT innovation platform database, obtain the basic information of the source data, and set data tags for each piece of data;
[0009] Process the source data, read the fields and contents in the migration data, and construct a data similarity verification model;
[0010] Select the migration path and the target database according to the migration requirements, and dynamically evaluate the similarity between the source data and the target data through the data similarity verification model;
[0011] When the data table reaches each node along the migration path, verify the migrated source data based on the target database, and perform incremental verification on the data of different fields in the table;
[0012] During the data migration process, real-time verification is performed on the data tables of the source data to be migrated between different migration nodes. Through the incremental verification mechanism, the data changes at each node are gradually verified to make the incremental data of each field consistent with the source data in the target database;
[0013] If data loss or inconsistent verification occurs during the verification process, trigger the alarm mechanism.
[0014] As a further solution of the present invention, select the database driver according to the database type used by the Xinchuang platform, use the database driver and the database connection string containing the database address and verification information to establish a connection with the database, execute the SQL query statement to read the source data to be migrated, and read and store the query results into the data structure in the program.
[0015] As a further solution of the present invention, obtaining the basic information of the source data includes querying the database metadata to obtain the table structure information. The table structure information includes the field name, data type, and constraint conditions; when setting the data label for each piece of data, the data label elements include the unique identifier field in the source data table, the data source identifier, the data type, and the data creation timestamp. A data object formed by combining the data label elements into a string forms a unique data label.
[0016] As a further solution of the present invention, process the source data, read the fields and contents in the migration data, including the following steps:
[0017] Read the field names and data types of the target data table;
[0018] Read the field names and data types in the table structure information of the source data;
[0019] Obtain the content of the target data table from the target database through the SQL query statement, and extract and convert the source data and the target data according to the field mapping relationship;
[0020] When the field names or formats do not match, use Java to map the source data fields to the target data fields.
[0021] As a further solution of the present invention, when constructing the data similarity verification model, it includes the following steps:
[0022] Determine the string similarity between the source data and the target data according to the data type, calculate the minimum edit distance between the two strings, and calculate the similarity of the text data;
[0023] Calculate the difference between the dates of the source data and the target data according to the data creation timestamp, and calculate the similarity using the absolute value of the date difference;
[0024] Comprehensively obtain the comprehensive similarity between the source data fields and the target data fields by weighting and integrating the numerical similarities of all fields and the date differences.
[0025] As a further solution of the present invention, selecting a migration path and a target database according to the migration requirements includes the following steps:
[0026] Obtain the amount of data to be migrated, the migration speed, and the type of the target database;
[0027] Determine a compatible target database according to the source data structure, the table structure of the target database, and the data type;
[0028] Determine the data migration path according to the migration requirements and the selection of the target database; where:
[0029] When the source database and the target database structures are completely compatible, determine the direct data migration path;
[0030] There are structural differences between the source database and the target database. First, migrate the data to the intermediate layer, perform data cleaning and conversion, and then migrate the data to the target database;
[0031] If the source data changes, migrate the incremental data through the incremental migration path to update the target database.
[0032] As a further solution of the present invention, dynamically evaluate the similarity between the source data and the target data through a data similarity verification model, and also include setting a similarity threshold. If the similarity is lower than the similarity threshold, trigger an alarm and mark the data as unmatched, where the set similarity threshold is a global threshold.
[0033] As a further solution of the present invention, dynamically evaluating the similarity between the source data and the target data through a data similarity verification model includes the following steps:
[0034] Initialize the field mapping relationship between the source data and the target data according to the table structure information and the field mapping rules;
[0035] According to the field types of the source data and the target data, calculate the similarity of each field using the minimum edit distance and the date difference;
[0036] The similarity of all fields is weighted and synthesized according to the field weights to obtain the comprehensive similarity of the entire data record;
[0037] According to the preset similarity threshold, determine whether the data migration is successful; if the comprehensive similarity is lower than the set threshold, re-evaluate the data mapping or perform manual intervention.
[0038] As a further solution of the present invention, when the data table reaches each node along the migration path, design the migration path, including the first, second, and third migration nodes. Each node is used to verify and back up the data, and establish connections between nodes; wherein, the migration path and nodes are defined as:
[0039] The first migration node: used to receive the source data, perform preliminary verification, and back up the data;
[0040] The second migration node: this node further performs verification, focusing on verifying the accuracy and consistency of the incremental data;
[0041] The third migration node: the final node, responsible for migrating the data to the target database and performing the last verification to ensure data consistency;
[0042] Among them, when establishing connections between nodes, connect the first migration nodes on the migration path to ensure that the source data is verified when it reaches the first migration node; connect the second migration nodes on the path to ensure that the migrated data is checked when it reaches the second node; connect the third migration nodes on the path to ensure that the final data is verified when it reaches the target database.
[0043] As a further solution of the present invention, during the data migration process, real-time verification is performed on the data table of the source data to be migrated between different migration nodes, including the following steps:
[0044] Verification when the data reaches the first migration node: At the first migration node, back up the source data and perform field-level verification. For each field of the data table, perform data integrity checks, including: verifying whether the field types and lengths in the source data table conform to the field specifications of the target data table; verifying whether the data conforms to the non-null constraint of the field, and if the field should be non-null, verify whether there are null values; verifying the consistency of the field data types, and whether numbers and dates can be successfully converted into the formats supported by the target database; if data loss or non-compliance is found at the first migration node, re-select the migration path;
[0045] Verification when the data reaches the second migration node: At the second migration node, perform incremental verification to check whether the newly added or modified data since the first verification is consistent with the target database; verify the incremental data according to the timestamp, log record or ID serial number of the incremental mark, and compare the differences between the source data and the target data;
[0046] Verification when data arrives at the third migration node: When the data arrives at the third migration node, perform a final verification to check whether the field data of the data table in the target database is consistent with the source database; use the data similarity verification model for verification. If data inconsistency or missing data is found at the third migration node, trigger the data recovery mechanism to repair the data, re-migrate the missing data, and repair the data with format errors.
[0047] As a further solution of the present invention, the data recovery mechanism includes:
[0048] Automatic rollback: If the incremental data verification fails, roll back to the previous migration node and re-select other paths for data migration;
[0049] Automatic retransmission: Through the incremental identifier, re-migrate the missing incremental data;
[0050] Alternative migration path: If one path fails, automatically switch to the alternative path.
[0051] In a second aspect, the present invention also provides a data migration verification system for an information innovation platform, including the following components:
[0052] Data acquisition module: Used to extract the data to be migrated from the source database of the information innovation platform, including reading the basic information, table structure information, and data field content of the source data, and setting a unique data label for each piece of data.
[0053] Data processing module: Used to process and transform the source data, read the table structure information of the target database, perform data field mapping, and process the matching of field types and formats according to preset rules to ensure that the source data can be correctly migrated to the target database.
[0054] Data similarity verification module: Calculate the similarity between the source data and the target data based on information such as data type, field content, and data creation timestamp, and dynamically evaluate whether the data is consistent. If the similarity is lower than the set threshold, trigger the alarm mechanism and mark the data as unmatched. This module also supports judging whether the data migration is successful based on similarity calculation and preset rules.
[0055] Migration path selection module: Automatically select the optimal migration path according to factors such as migration requirements, compatibility of the target database, and data volume. This module can determine whether to directly migrate, or whether data cleaning and transformation need to be performed through an intermediate layer, or an incremental migration path is used for data update.
[0056] Node Verification Module: Responsible for verifying data at each migration node. The first node performs preliminary verification and backs up the data; the second node performs incremental verification to ensure that the newly added or modified data is consistent with the target database; the third node performs final verification to ensure the consistency of the data in the target database. Each node can perform data integrity checks and ensure the accuracy and consistency of the data through field-level verification.
[0057] Alarm and Recovery Module: During the verification process, if data loss or verification inconsistency is found, the alarm mechanism is triggered to send notifications to relevant personnel and take corresponding measures according to the configured recovery mechanism. The recovery mechanism includes operations such as automatic rollback, incremental data retransmission, and standby migration path switching.
[0058] Data Backup and Recovery Module: Used to back up data during the migration process and perform recovery operations in case of data loss or verification inconsistency, ensuring that data is not lost or damaged during the migration process.
[0059] Data Migration Control Module: Used to control and schedule the entire data migration process, ensure the smooth migration of data from the source database to the target database, monitor the verification results of each node, perform dynamic adjustment of the migration path, and ensure the real-time and integrity of data migration.
[0060] Log Record and Tracking Module: Records various operations and verification results during the data migration process, including the verification situation of each migration node, data repair and rollback records, etc., for tracing problems and ensuring the transparency of data migration.
[0061] The data migration verification system of the Xinchuang platform of the present invention completes the data migration process of the Xinchuang platform accurately and efficiently through the cooperation of the above modules. At the same time, it provides multi-level data verification, repair, and alarm mechanisms, avoiding data loss or errors during the migration process, and ensuring the consistency and reliability of the migrated data.
[0062] Compared with the prior art, a data migration verification method and system for a Xinchuang platform proposed by the present invention has the following beneficial effects:
[0063] 1. The present invention dynamically evaluates the similarity between the source data and the target data through a data similarity verification model, and combines an incremental verification mechanism to ensure the accurate matching and consistency between the source data and the target data during the data migration process. The incremental data of each field is carefully verified. If inconsistencies or omissions occur, the alarm mechanism will be triggered to promptly identify and handle potential problems. Through the three-stage verification path (preliminary verification, incremental verification, and final verification), the present invention minimizes errors and data loss during the data migration process.
[0064] 2. The present invention intelligently selects a migration path according to factors such as migration requirements, data volume, and the type of target database. When the source database is compatible with the target database structure, data can be directly migrated; if there are structural differences, data cleaning and conversion are performed through an intermediate layer and then migrated. The introduction of the incremental migration path ensures flexibility during the data migration process, especially when the source data changes, and can effectively update the target database. This mechanism ensures the dynamic adjustment of the migration path and avoids data inconsistency problems during the migration process.
[0065] 3. The present invention adopts a field-level verification method based on data tags. Each piece of data is uniquely identified by means of data tags, and through the field mapping relationship between the source data and the target data, accurate matching of each field during the migration process is ensured. When the field name or format does not match, the system can automatically perform data conversion and adjustment, reducing the need for manual intervention and improving the automation level of the migration process.
[0066] 4. During the data migration process, a real-time verification mechanism is adopted to ensure that each migration node can be fully verified. If data is missing or the verification is inconsistent during the migration process, the system will immediately trigger an alarm mechanism and mark the relevant data as mismatched. In addition, the system also has a data recovery mechanism that can automatically roll back to the previous migration node when an exception occurs, reselect the migration path, and ensure the high reliability of data migration.
[0067] 5. The present invention provides a variety of data recovery means, including automatic rollback, incremental data retransmission, and standby migration path switching, etc., to ensure that when a failure occurs during the migration process, measures can be taken in a timely manner to repair it, minimizing the risk of data loss and migration failure. These mechanisms provide strong fault tolerance for data migration. Through the optimization of the data structure and intelligent data processing flow, the migration efficiency is improved. The source data undergoes comprehensive preprocessing and cleaning before migration, including field mapping, data type conversion, and timestamp verification, etc., to ensure that the source data can be quickly adapted to the target database. At the same time, the system can gradually verify the data tables, avoiding excessive performance consumption during full-scale migration, and improving the migration speed and resource utilization efficiency.
[0068] In summary, the method and system for verifying data migration in the domestic information technology innovation platform provided by the present invention greatly reduce the need for manual intervention through a highly automated verification and data processing mechanism. Especially when there are data format mismatches or changes in incremental data, the system can automatically perform operations such as field mapping and data conversion, and judge data consistency through a similarity verification model to ensure high-precision and high-efficiency data migration; through a multi-level and multi-dimensional verification mechanism, intelligent migration path selection, real-time verification and alarm, powerful recovery capabilities, and efficient data processing and optimization, the success rate, accuracy, and efficiency of data migration are greatly improved. At the same time, the cost of manual intervention is reduced, and the automation and security of the migration process are enhanced, which has important practical value.
[0069] These aspects or other aspects of the present application will be more clearly understood in the following description of the embodiments. It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] To more clearly illustrate the technical solutions in the embodiments of the present invention or related technologies, the following briefly introduces the drawings required for describing the exemplary embodiments or related technologies. The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the drawings:
[0071] Figure 1 It is a flowchart of a method for verifying data migration in a domestic information technology innovation platform according to an embodiment of the present invention.
[0072] Figure 2 It is a flowchart of reading fields and contents in migration data in a method for verifying data migration in a domestic information technology innovation platform according to an embodiment of the present invention.
[0073] Figure 3 It is a flowchart of constructing a data similarity verification model in a method for verifying data migration in a domestic information technology innovation platform according to an embodiment of the present invention.
[0074] Figure 4 It is a flowchart of selecting a migration path and a target database according to migration requirements in a method for verifying data migration in a domestic information technology innovation platform according to an embodiment of the present invention.
[0075] Figure 5 It is a flowchart of selecting a migration path and a target database according to migration requirements in a method for verifying data migration in a domestic information technology innovation platform according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0076] Next, in combination with the accompanying drawings and specific embodiments, the present application will be further described. It should be noted that, on the premise of no conflict, the following-described embodiments or technical features can be arbitrarily combined with each other to form new embodiments.
[0077] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the following further elaborates on the embodiments of the present invention in detail with reference to specific embodiments and the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0078] It should be noted that all the expressions using "first" and "second" in the embodiments of the present invention are used to distinguish two non-identical entities or non-identical parameters with the same name. It can be seen that "first" and "second" are only for the convenience of expression and should not be construed as a limitation on the embodiments of the present invention. In addition, the terms "include" and "have" and any of their variations are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units inherently includes other steps or units.
[0079] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the scope of protection of the present application.
[0080] The flowcharts shown in the accompanying drawings are only illustrative examples and do not necessarily include all the contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can be decomposed, combined, or partially merged, so the actual execution order may change according to the actual situation.
[0081] Next, in combination with the accompanying drawings, some embodiments of the present application will be described in detail. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0082] In the existing database migration, it not only involves the transmission of data itself, but also involves the verification of multiple dimensions such as data tags, migration tables, data source consistency, and similarity analysis. These traditional methods often ignore the relationships between data in different dimensions, upstream and downstream dependencies, and dynamic data changes during the migration process, resulting in the inability to detect data integrity and consistency problems in a timely manner, and these problems often only emerge after the migration is completed. The present invention proposes a data migration verification method and system for a domestic information and communication technology (ICT) innovation platform, which can, during the data migration process, verify and dynamically monitor multi-dimensional features such as the tags, basic information, and similarity analysis of the migrated data, and automatically identify and correct data missing or inconsistent problems during the real-time data migration process. This method ensures a high degree of reliability and consistency during the data migration process by introducing a similarity matching algorithm, dynamic data table verification, and comparison with target source data.
[0083] See Figure 1 As shown, an embodiment of the present invention provides a data migration verification method for a domestic ICT innovation platform, and this method includes the following steps:
[0084] Step S10: Read all source data to be migrated in the database of the domestic ICT innovation platform, obtain the basic information of the source data, and set data tags for each piece of data.
[0085] Step S20: Process the source data, read the fields and contents in the migrated data, and construct a data similarity verification model.
[0086] Step S30: Select a migration path and a target database according to the migration requirements, and dynamically evaluate the similarity between the source data and the target data through the data similarity verification model.
[0087] Step S40: When the data table arrives at each node along the migration path, verify the migrated source data based on the target database, and perform incremental verification on the data in different fields of the table.
[0088] Step S50: During the data migration process, perform real-time verification on the data tables of the source data to be migrated between different migration nodes, and gradually verify the data changes at each node through the incremental verification mechanism, so that the incremental data of each field is consistent with the source data in the target database.
[0089] Step S60: If data missing or verification inconsistency occurs during the verification process, trigger an alarm mechanism.
[0090] In this embodiment, the database driver is selected according to the database type used by the Xinchuang platform. The database driver and the database connection string containing the database address and authentication information are used to establish a connection with the database. The SQL query statement is executed to read the source data to be migrated, and the query results are read and stored in the data structure in the program. Exemplarily, the database type used by the Xinchuang platform can be MySQL, PostgreSQL, Oracle, or SQL Server, and the corresponding database driver is selected. Among them, for MySQL, MySQL Connector / J can be used; for PostgreSQL, pgjdbc can be used; for Oracle, Oracle JDBC Driver can be used; for SQL Server, Microsoft JDBC Driver can be used.
[0091] Among them, obtaining the basic information of the source data includes querying the database metadata to obtain the table structure information, and the table structure information includes field names, data types, and constraint conditions. When setting data tags for each piece of data, the data tag elements include the unique identifier field in the source data table, the data source identifier, the data type, and the data creation timestamp. A data object in the form of a string is formed by combining the data tag elements to form a unique data tag. Exemplarily, the tag format can be set in the format of ID - timestamp - source - status, such as 12345 - 2024 - 12 - 18T10:15:00 - sourceSystem - to be migrated.
[0092] In this embodiment, as shown in Figure 2 In step S20, the source data is processed, and the fields and contents in the migration data are read, including the following steps:
[0093] Step S201: Read the field names and data types of the target data table.
[0094] In this step, during the migration process, first, the structure information of the target data table in the target database needs to be read, including field names, data types, etc. This operation is usually implemented through database metadata queries, and the field names and data types of the target table can be obtained through SQL query statements.
[0095] Step S202: Read the field names and data types in the table structure information of the source data.
[0096] In this step, the structure information of the source data table in the source database is read to ensure that the types of the source data fields can be correctly understood. Similarly, SQL query statements are used to obtain the field information of the source table.
[0097] Step S203: Obtain the content of the target data table from the target database through an SQL query statement, and extract and transform the source data and the target data according to the field mapping relationship.
[0098] In this step, use an SQL query to extract the specific content of the target data table from the target database for subsequent comparison with the source data. For example, obtain all the data in the target table through an SQL statement and save it in the form of a data frame, dictionary, etc.
[0099] Step S204: When the field names or formats do not match, use Java to map the source data fields to the target data fields.
[0100] During the migration process of this step, if there is a situation where the field names or data types between the source data and the target data table do not match, perform field mapping to ensure that the source data can be correctly migrated to the target table. Field mapping can be implemented using Java code. For example, use a mapping table or program logic to map the fields of the source data table to the fields of the target data table one by one.
[0101] In this embodiment, as shown in Figure 3 When constructing the data similarity verification model in step S20, the following steps are included:
[0102] Step S211: Determine the string similarity between the source data and the target data according to the data type, calculate the minimum edit distance between the two strings, and calculate the similarity of the text data.
[0103] In this step, for fields of text type, the matching degree between the source data and the target data can be judged by calculating the string similarity between them. Calculate using the minimum edit distance (Levenshtein Distance).
[0104] Step S212: Calculate the difference between the dates of the source data and the target data according to the data creation timestamp, and calculate the similarity using the absolute value of the date difference;
[0105] Step S213: Comprehensively obtain the comprehensive similarity between the source data fields and the target data fields by weighting the numerical similarity and date difference of all fields.
[0106] Weight the similarity values of each field (including string similarity and date difference) according to a certain weight, and finally obtain the comprehensive similarity. According to the preset weights (for example, the weight of text fields is 0.7 and the weight of date fields is 0.3), weight and average the similarity values of each field to obtain the overall similarity.
[0107] In this embodiment, as shown in Figure 4As shown, in step S30, selecting the migration path and the target database according to the migration requirements includes the following steps:
[0108] Step S301, obtain the amount of data to be migrated, the migration speed, and the type of the target database;
[0109] Step S302, determine the target database with compatibility according to the source data structure, the table structure of the target database, and the data type;
[0110] Step S303, determine the data migration path according to the migration requirements and the selection of the target database; where: if the source data is compatible with the target database structure, it can be directly migrated. If there are structural differences, it can be first migrated to the intermediate layer for cleaning and conversion, and then migrated to the target database. If the source data will change, an incremental migration path can be adopted.
[0111] Step S304, when the source database and the target database structures are completely compatible, determine the path for direct data migration;
[0112] Step S305, when there are structural differences between the source database and the target database, first migrate the data to the intermediate layer, perform data cleaning and conversion, and then migrate the data to the target database.
[0113] Among them, when there are structural differences between the source data and the target database, first migrate the data to the intermediate layer for data cleaning and conversion, and then migrate it to the target database. Process the data in the intermediate layer: convert the data type, handle missing fields, rename fields, etc., and finally migrate the cleaned data to the target database.
[0114] Step S306, if the source data changes, migrate the incremental data through the incremental migration path to update the target database.
[0115] Among them, if the data in the source database changes (such as new additions or modifications), it is necessary to migrate the changed data through the incremental migration path. According to the timestamp or incremental identifier (such as log files, ID serial numbers, etc.), obtain the newly added or modified records, and migrate these incremental data to the target database. Incremental migration is usually carried out through scheduled tasks or triggers.
[0116] The present invention intelligently selects the migration path according to factors such as migration requirements, data volume, and the type of the target database. When the source database is compatible with the target database structure, data can be directly migrated; if there are structural differences, data cleaning and conversion are performed through the intermediate layer and then migrated. The introduction of the incremental migration path ensures the flexibility in the data migration process, especially when the source data changes, it can effectively update the target database. This mechanism ensures the dynamic adjustment of the migration path and avoids data inconsistency problems during the migration process.
[0117] Among them, the similarity between the source data and the target data is dynamically evaluated through a data similarity verification model, and it also includes setting a similarity threshold. If the similarity is lower than the similarity threshold, an alarm is triggered and the data is marked as unmatched. The set similarity threshold is a global threshold.
[0118] In this embodiment, as shown in Figure 5 the following steps are included in dynamically evaluating the similarity between the source data and the target data through a data similarity verification model:
[0119] Step S311: Initialize the field mapping relationship between the source data and the target data according to the table structure information and the field mapping rules;
[0120] Step S312: Calculate the similarity of each field using the minimum edit distance and date difference according to the field types of the source data and the target data.
[0121] In this step, for each field, the similarity between the source data and the target data is calculated. For string-type fields, the edit distance is used, and for date-type fields, the time difference is used. Among them, the Levenshtein distance algorithm is used to calculate the string similarity, and the date difference is used to calculate the similarity of date-type fields.
[0122] Step S313: Weight and synthesize the similarities of all fields according to the field weights to obtain the comprehensive similarity of the entire data record.
[0123] Among them, the similarity values are weighted according to the importance of different fields to calculate the comprehensive similarity of the entire data record. By defining the field weights, the similarity values of each field are weighted and synthesized according to the weights to obtain the comprehensive similarity.
[0124] Step S314: Judge whether the data migration is successful according to the preset similarity threshold; if the comprehensive similarity is lower than the set threshold, re-evaluate the data mapping or perform manual intervention.
[0125] Among them, it is judged whether the data migration is successful according to the comprehensive similarity and the set threshold. If the similarity is lower than the set threshold, it is considered that the migration fails and manual intervention or re-mapping of fields is required.
[0126] The present invention adopts a field-level verification method based on data tags. Each piece of data is uniquely identified by means of data tags, and through the field mapping relationship between the source data and the target data, accurate matching of each field during the migration process is ensured. When the field names or formats do not match, the system can automatically perform data conversion and adjustment, reducing the need for manual intervention and improving the automation degree of the migration process.
[0127] Among them, when the data table reaches each node along the migration path, a migration path is designed, including the first, second, and third migration nodes. Each node is used to verify and back up the data, and connections between nodes are established; among them, the migration path and nodes are defined as:
[0128] The first migration node: used to receive the source data, perform preliminary verification, and back up the data;
[0129] The second migration node: this node further performs verification, focusing on verifying the accuracy and consistency of the incremental data;
[0130] The third migration node: the final node, responsible for migrating the data to the target database and performing the last verification to ensure data consistency;
[0131] Among them, when establishing connections between nodes, the first migration nodes on the migration path are connected to ensure that the source data is verified when it reaches the first migration node; the second migration nodes are connected on the path to ensure that the migrated data is checked when it reaches the second node; the third migration nodes are connected on the path to ensure that the final data is verified when it reaches the target database.
[0132] In this embodiment, during the data migration process, real-time verification is performed on the data table of the source data to be migrated between different migration nodes, including the following steps:
[0133] Verification when the data reaches the first migration node: At the first migration node, the source data is backed up, and field-level verification is performed. For each field of the data table, data integrity checks are carried out, including: verifying whether the field types and lengths in the source data table conform to the field specifications of the target data table; verifying whether the data conforms to the non-null constraint of the field, and if the field should be non-null, verifying whether there are null values; verifying the consistency of the field data types, whether numbers and dates can be successfully converted into the formats supported by the target database; if data is found to be missing or non-compliant at the first migration node, a new migration path is reselected;
[0134] Verification when the data reaches the second migration node: At the second migration node, incremental verification is performed to check whether the newly added or modified data since the first verification is consistent with the target database; the incremental data is verified according to the timestamp, log record or ID serial number of the incremental marker, and the differences between the source data and the target data are compared;
[0135] Verification when the data reaches the third migration node: When the data reaches the third migration node, final verification is performed to verify whether the field data in the data table in the target database is consistent with the source database; a data similarity verification model is used for verification. If data inconsistency or missing is found at the third migration node, a data recovery mechanism is triggered to perform data repair, re-migrate the missing data, and repair the data with format errors.
[0136] During the data migration process, a real-time verification mechanism is adopted to ensure that each migration node can conduct sufficient verification. If data loss or verification inconsistency occurs during the migration, the system will immediately trigger the alarm mechanism and mark the relevant data as mismatched. In addition, the system also has a data recovery mechanism that can automatically roll back to the previous migration node in case of anomalies, reselect the migration path, and ensure the high reliability of data migration.
[0137] Among them, the data recovery mechanism includes:
[0138] Automatic rollback: If the incremental data verification fails, roll back to the previous migration node and reselect other paths for data migration;
[0139] Automatic retransmission: Through the incremental identifier, re-migrate the missing incremental data;
[0140] Standby migration path: If one path fails, automatically switch to the standby path.
[0141] The present invention provides a variety of data recovery means, including automatic rollback, incremental data retransmission, and standby migration path switching, etc., to ensure that when a failure occurs during the migration process, measures can be taken in a timely manner for repair, minimizing the risk of data loss and migration failure to the greatest extent. These mechanisms provide strong fault tolerance for data migration. Through the optimization of the data structure and intelligent data processing processes, the migration efficiency is improved. The source data undergoes comprehensive preprocessing and cleaning before migration, including field mapping, data type conversion, and timestamp verification, etc., to ensure that the source data can be quickly adapted to the target database. At the same time, the system can gradually verify the data table, avoiding excessive performance consumption during full-scale migration, and improving the migration speed and resource utilization efficiency.
[0142] The present invention dynamically evaluates the similarity between the source data and the target data through a data similarity verification model, and combines an incremental verification mechanism to ensure the accurate matching and consistency between the source data and the target data during the data migration process. The incremental data of each field is carefully verified. If inconsistency or loss occurs, the alarm mechanism will be triggered to identify and handle potential problems in a timely manner. Through the three-stage verification path (preliminary verification, incremental verification, and final verification), the present invention minimizes errors and data loss during the data migration process.
[0143] In summary, the data migration verification method and system provided by the present invention greatly reduce the need for manual intervention through a highly automated verification and data processing mechanism. Especially when there are data format mismatches or changes in incremental data, the system can automatically perform operations such as field mapping and data conversion, and judge data consistency through a similarity verification model to ensure high-precision and high-efficiency data migration; through a multi-level and multi-dimensional verification mechanism, intelligent migration path selection, real-time verification and alarm, powerful recovery capabilities, as well as efficient data processing and optimization, it greatly improves the success rate, accuracy and efficiency of data migration, while reducing the cost of human intervention, enhancing the automation and security of the migration process, and having important practical value.
[0144] It should be noted that the above-mentioned drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present invention, rather than for limiting purposes. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the time sequence of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously in, for example, multiple modules.
[0145] It should be understood that although the above is described in a certain order, these steps are not necessarily executed in the above order successively. Unless there is a clear indication in this article, the execution of these steps has no strict order restriction, and these steps can be executed in other orders. Moreover, a part of the steps in this embodiment may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same moment, but can be executed at different moments, and the execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0146] In the second aspect of the embodiments of the present invention, the present invention also provides a data migration verification system for an information and communication technology (ICT) innovation platform, including:
[0147] A data collection module: used to extract the data to be migrated from the source database of the ICT innovation platform, including reading the basic information, table structure information and data field content of the source data, and setting a unique data label for each piece of data.
[0148] A data processing module: used to process and convert the source data, read the table structure information of the target database, perform data field mapping, and process the matching of field types and formats according to preset rules to ensure that the source data can be correctly migrated to the target database.
[0149] Data Similarity Verification Module: Calculate the similarity between the source data and the target data based on information such as data type, field content, data creation timestamp, etc., and dynamically evaluate whether the data is consistent. If the similarity is lower than the set threshold, trigger the alarm mechanism and mark the data as mismatched. This module also supports judging whether the data migration is successful according to the similarity calculation and preset rules.
[0150] Migration Path Selection Module: Automatically select the optimal migration path according to factors such as migration requirements, compatibility of the target database, and data volume. This module can determine whether to perform a direct migration, or whether data cleaning and conversion need to be carried out through an intermediate layer, or an incremental migration path is adopted for data update.
[0151] Node Verification Module: Responsible for verifying the data at each migration node. The first node performs a preliminary verification and backs up the data; the second node performs an incremental verification to ensure that the newly added or modified data is consistent with the target database; the third node performs a final verification to ensure the consistency of the data in the target database. Each node can perform data integrity checks and ensure the accuracy and consistency of the data through field-level verification.
[0152] Alarm and Recovery Module: During the verification process, if data loss or verification inconsistency is found, trigger the alarm mechanism, send notifications to relevant personnel, and take corresponding measures according to the configured recovery mechanism. The recovery mechanism includes operations such as automatic rollback, incremental data retransmission, and standby migration path switching.
[0153] Data Backup and Recovery Module: Used to back up data during the migration process and perform recovery operations in case of data loss or verification inconsistency, ensuring that data is not lost or damaged during the migration process.
[0154] Data Migration Control Module: Used to control and schedule the entire data migration process, ensure the smooth migration of data from the source database to the target database, monitor the verification results of each node, perform dynamic adjustment of the migration path, and ensure the real-time and integrity of data migration.
[0155] Log Recording and Tracking Module: Record various operations and verification results during the data migration process, including the verification situation of each migration node, data repair and rollback records, etc., in order to trace problems and ensure the transparency of data migration.
[0156] The data migration verification system of the Xinchuang platform of the present invention ensures that the data migration process of the Xinchuang platform can be completed accurately and efficiently through the cooperation of the above modules, and at the same time provides multi-level data verification, repair and alarm mechanisms, avoiding data loss or errors during the migration process, and ensuring the consistency and reliability of the migrated data.
[0157] Through the above detailed steps, the data migration verification system of the domestic innovation platform of the present invention is used to execute the steps of the data migration verification method of the domestic innovation platform in the above embodiments, which will not be elaborated here.
[0158] The above are exemplary embodiments disclosed by the present invention. However, it should be noted that various changes and modifications can be made without departing from the scope of the embodiments disclosed by the present invention as defined by the claims. The functions, steps, and / or actions of the method claims according to the disclosed embodiments herein need not be performed in any particular order. In addition, although the elements disclosed by the embodiments of the present invention can be described or claimed in an individual form, they can also be understood as plural unless explicitly limited to the singular.
[0159] It should be understood that, as used herein, unless the context clearly supports an exception, the singular form "a" is also intended to include the plural form. It should also be understood that the "and / or" used herein refers to any and all possible combinations of one or more of the related listed items. The serial numbers of the disclosed embodiments of the present invention above are only for description and do not represent the superiority or inferiority of the embodiments.
[0160] Those of ordinary skill in the art should understand that: the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the embodiments disclosed by the present invention (including the claims) is limited to these examples; under the concept of the embodiments of the present invention, the technical features between the above embodiments or different embodiments can also be combined, and there are many other variations in different aspects of the embodiments of the present invention as above, which are not provided in detail for the sake of brevity. Therefore, any omission, modification, equivalent replacement, improvement, etc. made within the spirit and principle of the embodiments of the present invention shall be included in the protection scope of the embodiments of the present invention.
Claims
1. A data migration verification method for an information innovation platform, characterized in that: The method comprises the following steps: Read all source data to be migrated in the Xinchuang platform database, obtain basic information about the source data, and set data labels for each piece of data; Process the source data, read the fields and content in the migrated data, and build a data similarity verification model; Select the migration path and target database based on the migration requirements, and dynamically evaluate the similarity between the source data and the target data through the data similarity verification model; When the data table reaches each node along the migration path, the migrated source data is verified based on the target database, and the data of different fields in the table are incrementally verified; During the data migration process, the data tables of the source data to be migrated are verified in real time between different migration nodes. The data changes on each node are gradually verified through the incremental verification mechanism, so that the incremental data of each field in the target database is consistent with the source data; If data is missing or verification is inconsistent during the verification process, the alarm mechanism will be triggered; Among them, when building a data similarity verification model, the following steps are included: Determine the string similarity between the source data and the target data according to the data type, calculate the minimum edit distance between the two strings, and calculate the similarity of the text data; Calculate the difference between the source data and target data dates based on the data creation timestamp, and use the absolute value of the date difference to calculate the similarity; The numerical similarity and date difference of all fields are weighted and integrated according to the fields to obtain the comprehensive similarity between the source data field and the target data field; The similarity between the source data and the target data is dynamically evaluated by a data similarity verification model, including the following steps: Initialize the field mapping relationship between source data and target data according to the table structure information and field mapping rules; Based on the field types of the source and target data, the similarity of each field is calculated using the minimum edit distance and date difference; The similarities of all fields are weighted and synthesized according to the field weights to obtain the comprehensive similarity of the entire data record; Determine whether the data migration is successful based on the preset similarity threshold; if the comprehensive similarity is lower than the set threshold, re-evaluate the data mapping or perform manual intervention.
2. The data migration verification method of the Xinchuang platform as described in claim 1 is characterized in that: Select a database driver according to the database type used by the trusted computing platform, use the database driver and the database connection string containing the database address and verification information to establish a connection with the database, execute SQL query statements to read the source data to be migrated, and read and store the query results into the data structure in the program.
3. The data migration verification method of the Xinchuang platform as described in claim 2 is characterized in that: Obtaining basic information about source data includes querying database metadata to obtain table structure information, which includes field names, data types, and constraints. When setting data labels for each piece of data, data label elements include a unique identifier field in the source data table, data source identifier, data type, and data creation timestamp. Data objects that are combined into strings through data label elements form unique data labels.
4. The data migration verification method of the Xinchuang platform as described in claim 3 is characterized in that: Process the source data and read the fields and contents in the migration data, including the following steps: Read the field name and data type of the target data table; Read the field name and data type in the table structure information of the source data; Obtain the contents of the target data table from the target database through SQL query statements, and extract and convert the source data and target data according to the field mapping relationship; Use Java to map source data fields to target data fields when field names or formats do not match.
5. The data migration verification method of the Xinchuang platform as described in claim 3 is characterized in that: Selecting the migration path and target database based on migration requirements includes the following steps: Get the amount of data to be migrated, the migration speed, and the type of target database; Determine a compatible target database based on the source data structure and the table structure and data type of the target database; Determine the data migration path based on the migration requirements and the target database selection; among which: When the source database and target database structures are fully compatible, determine the path for direct data migration; If there are structural differences between the source database and the target database, first migrate the data to the middle layer, clean and convert the data, and then migrate the data to the target database; If the source data changes, the incremental data is migrated through the incremental migration path to update the target database.
6. The data migration verification method of the Xinchuang platform as described in claim 5 is characterized in that: The similarity between the source data and the target data is dynamically evaluated through the data similarity verification model, which also includes setting a similarity threshold. If the similarity is lower than the similarity threshold, an alarm is triggered and the data is marked as mismatched, wherein the set similarity threshold is a global threshold.
7. The data migration verification method of the Xinchuang platform as described in claim 6 is characterized in that: When the data table reaches each node along the migration path, a migration path is designed, including the first, second, and third migration nodes. Each node is used to verify and back up the data and establish a connection between nodes. The migration path and nodes are defined as follows: The first migration node is used to receive source data, perform preliminary verification, and perform data backup; Second migration node: This node is further verified, focusing on the accuracy and consistency of incremental data; The third migration node: the final node, responsible for migrating data to the target database and performing the final verification to ensure data consistency; Among them, when establishing the connection between nodes, the first migration node on the migration path is connected to ensure that the source data is verified at the first migration node; the second migration node is connected on the path to ensure that the migration data is checked at the second node; the third migration node is connected on the path to ensure that the final data is verified at the target database.
8. A data migration verification system for an information innovation platform, characterized in that: Used to execute the Xinchuang platform data migration verification method as claimed in claim 7, the system includes: Data collection module: used to extract the data to be migrated from the source database of the Xinchuang platform, including reading the basic information, table structure information and data field content of the source data, and setting a unique data label for each piece of data; Data processing module: used to process and convert source data, read the table structure information of the target database, perform data field mapping, and process the matching of field types and formats according to preset rules; Data similarity verification module: Calculates the similarity between source data and target data based on data type, field content, and data creation timestamp information, and dynamically evaluates whether the data is consistent; if the similarity is lower than the set threshold, an alarm mechanism is triggered and the data is marked as mismatched; Migration path selection module: automatically selects the optimal migration path based on migration requirements, compatibility of the target database, and data volume; Node verification module: responsible for verifying data at each migration node; Alarm and recovery module: During the verification process, if data is missing or the verification is inconsistent, the alarm mechanism will be triggered.
Citation Information
Patent Citations
Data migration method and device, computer readable storage medium and electronic equipment
CN118733566A