Data migration method, device, electronic device and storage medium
By obtaining and using database configuration information, converting source data into Flink two-dimensional data intermediate format and converting it into target database format, it solves the problem of difficult and inefficient data migration between different databases, and realizes efficient and secure data migration.
Patent Information
- Application Number
- CN202210045060.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-14
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-01-14
AI Technical Summary
Due to the inconsistent data structures between different databases, data migration between different databases is difficult and low in migration efficiency, which affects data security and database management costs during data migration.
By obtaining migration configuration information, the source data is obtained from the source database based on the configuration information of the source database and saved it as Flink-based two-dimensional data intermediate data. Then, based on the configuration information of the target database, the intermediate data is converted into the target data in the data format of the target database, and the target data is verified. If the verification result is normal, the target data is stored in the target database.
Data migration between different databases is realized, migration efficiency is improved, data security is enhanced, and database management costs are reduced.
Smart Images

Figure CN114372043B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of big data and data storage, and in particular to a data migration method, device, electronic device and storage medium. Background Art
[0002] With the rise and widespread application of big data architecture, enterprises are also gradually promoting the migration of traditional databases to big data architecture. At present, a situation has been formed where big data architecture databases and traditional databases coexist. The application of big data architecture solves the expansion capabilities of traditional databases in terms of distributed expansion, data storage, computing time, etc., but it also brings about problems such as complex architecture, diverse technology stacks, and inconsistent data structures.
[0003] In the existing technology, for database systems and tools such as hive, hbase, hdfs, and redis under the big data architecture, data is usually read, written, and managed based on one of them to meet specific business needs. With the expansion of business and the continuous increase in data scale, it is necessary to merge and migrate data between different databases.
[0004] However, due to the inconsistent data structures between different databases, data migration between different databases is difficult and the migration efficiency is low, which affects the data security and database management costs during the data migration process. Summary of the invention
[0005] The present application provides a data migration method, device, electronic device and storage medium to solve the problem of difficulty and low migration efficiency in data migration between different databases due to the inconsistency of data structures between different databases.
[0006] In a first aspect, the present application provides a data migration method, comprising:
[0007] Obtain migration configuration information, the migration configuration information including configuration information of a first type of source database and configuration information of a second type of target database, the configuration information being used to perform read / write operations on data in databases of corresponding types; based on the configuration information of the source database, obtain source data from the source database, and save the source data as intermediate data, wherein the intermediate data is two-dimensional data based on Flink; based on the configuration information of the target database, convert the intermediate data into target data in a data format of the target database, verify the target data, and obtain a verification result; if the verification result is normal, store the target data in the target database.
[0008] In a possible implementation, the source data includes multiple source data units, and based on the configuration information of the source database, the source data is obtained from the source database, and the source data is saved as intermediate data, including: obtaining a preset data export object, the data export object includes export parameters, and the data export object is used to stream read each source data unit from the database indicated by the export parameters; based on the configuration information of the source database, the export parameters of the data export object are obtained, and the data export object is called to stream read each source data unit in the source database; the obtained source data units are written into the memory in sequence in the format of the two-dimensional data to obtain the intermediate data.
[0009] In a possible implementation, the method further includes: after each reading of a source data unit in the source database, detecting a reading result corresponding to a current source data unit, the reading result indicating whether the source data unit has been read successfully; based on the reading result, if the current source data unit has been read successfully, recording a data identifier corresponding to the current source data unit; if the source data unit has not been read successfully, re-calling the data reading object to read the current source data unit based on a verification identifier corresponding to a previous source data unit.
[0010] In a possible implementation, the intermediate data includes multiple intermediate data units, and the conversion of the intermediate data into target data in the data format of the target database based on the configuration information of the target database includes: obtaining a preset data import object, the data import object including import parameters, and the data import object is used to convert the intermediate data unit into target data matching the data format of the database indicated by the import parameters; based on the configuration information of the target database, obtaining the import parameters of the data import object, and calling the data import object to convert the intermediate data unit into the target data.
[0011] In a possible implementation, the source data includes multiple source data units, the target data includes multiple target data units, and the target data is verified, including: obtaining a first number corresponding to the source data units, and a second number of the target data units; and obtaining a verification result based on the first number and the second number.
[0012] In a possible implementation, a verification result is obtained based on the first quantity and the second quantity, including: when the first quantity is equal to the second quantity, obtaining first sampling data and second sampling data, wherein the first sampling data includes multiple source data units, and the second sampling data includes multiple target data units of the same number, and the data identifier of each source data unit corresponds to the same data identifier of each target data unit; and performing a verification based on the MD5 value of the first sampling data and the MD5 value of the second sampling data to obtain a verification result, wherein if the MD5 value of the first sampling data is the same as the MD5 value of the second sampling data, the verification result is normal.
[0013] In a possible implementation manner, the configuration information includes a first information item and a second information item, wherein the first information item represents a database type of a corresponding database, and the second information item represents database connection information of the corresponding database.
[0014] In a second aspect, the present application provides a data migration device, including:
[0015] A configuration module, used to obtain migration configuration information, wherein the migration configuration information includes configuration information of a first type of source database and configuration information of a second type of target database, wherein the configuration information is used to perform read / write operations on data in databases of corresponding types;
[0016] A reading module, used for acquiring source data from the source database based on the configuration information of the source database, and saving the source data as intermediate data, wherein the intermediate data is two-dimensional data based on Flink;
[0017] The writing module is used to convert the intermediate data into target data in the data format of the target database based on the configuration information of the target database, and verify the target data to obtain a verification result. If the verification result is normal, the target data is stored in the target database.
[0018] In a possible implementation, the source data includes multiple source data units, and the reading module is specifically used to: obtain a preset data export object, the data export object includes export parameters, and the data export object is used to stream read each source data unit from the database indicated by the export parameters; based on the configuration information of the source database, obtain the export parameters of the data export object, and call the data export object to stream read each source data unit in the source database; write the obtained source data units into the memory in sequence in the format of the two-dimensional data to obtain the intermediate data.
[0019] In a possible implementation, the reading module is further used to: detect a reading result corresponding to a current source data unit after each reading of a source data unit in the source database, wherein the reading result indicates whether the source data unit is read successfully; based on the reading result, if the current source data unit is read successfully, record a data identifier corresponding to the current source data unit; if the source data unit is not read successfully, re-call the data reading object to read the current source data unit based on a verification identifier corresponding to a previous source data unit.
[0020] In a possible implementation, the intermediate data includes multiple intermediate data units. When the writing module converts the intermediate data into target data in the data format of the target database based on the configuration information of the target database, it is specifically used to: obtain a preset data import object, the data import object includes import parameters, and the data import object is used to convert the intermediate data unit into target data matching the data format of the database indicated by the import parameters; based on the configuration information of the target database, obtain the import parameters of the data import object, and call the data import object to convert the intermediate data unit into the target data.
[0021] In one possible implementation, the source data includes multiple source data units, and the target data includes multiple target data units. When verifying the target data, the write module is specifically used to: obtain a first quantity corresponding to the source data units, and a second quantity of the target data units; and obtain a verification result based on the first quantity and the second quantity, wherein if the first quantity is the same as the second quantity, the verification result is normal.
[0022] In a possible implementation, when the writing module obtains the verification result based on the first quantity and the second quantity, it is specifically used to: when the first quantity is equal to the second quantity, obtain first sampling data and second sampling data, wherein the first sampling data includes multiple source data units, and the second sampling data includes multiple target data units of the same number, and the data identifier of each source data unit corresponds to the same data identifier of each target data unit; perform verification based on the MD5 value of the first sampling data and the MD5 value of the second sampling data to obtain a verification result, wherein if the MD5 value of the first sampling data is the same as the MD5 value of the second sampling data, the verification result is normal.
[0023] In a possible implementation manner, the configuration information includes a first information item and a second information item, wherein the first information item represents a database type of a corresponding database, and the second information item represents database connection information of the corresponding database.
[0024] In a third aspect, the present application provides an electronic device, comprising: a processor, and a memory communicatively connected to the processor;
[0025] The memory stores computer-executable instructions;
[0026] The processor executes the computer-executable instructions stored in the memory to implement the data migration method as described in any one of the first aspects of the embodiments of the present application.
[0027] In a fourth aspect, the present application provides a computer-readable storage medium, in which computer execution instructions are stored. When the computer execution instructions are executed by a processor, they are used to implement the data migration method as described in any one of the first aspects of the embodiments of the present application.
[0028] According to a fifth aspect of an embodiment of the present application, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the data migration method as described in any one of the first aspects above.
[0029] The data migration method, device, electronic device and storage medium provided by the present application obtain migration configuration information, wherein the migration configuration information includes configuration information of a first type of source database and configuration information of a second type of target database, and the configuration information is used to perform read / write operations on data in databases of corresponding types; based on the configuration information of the source database, source data is obtained from the source database, and the source data is saved as intermediate data, wherein the intermediate data is two-dimensional data based on Flink; based on the configuration information of the target database, the intermediate data is converted into target data in the data format of the target database, and the target data is verified to obtain a verification result, and if the verification result is normal, the target data is stored in the target database. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0031] Figure 1 An application scenario diagram of the data migration method provided in an embodiment of the present application;
[0032] Figure 2 A flowchart of a data migration method provided by an embodiment of the present application;
[0033] Figure 3 A flowchart of a data migration method provided by another embodiment of the present application;
[0034] Figure 4 for Figure 3 A flowchart of streaming reading of each source data unit in the source database in step S203 in the illustrated embodiment;
[0035] Figure 5 for Figure 3 A flowchart of converting the intermediate data unit into the target data in step S206 in the illustrated embodiment;
[0036] Figure 6 A data verification process is provided in accordance with an embodiment of the present application;
[0037] Figure 7 A schematic diagram of the structure of a data migration device provided in one embodiment of the present application;
[0038] Figure 8 A schematic diagram of an electronic device provided by an embodiment of the present application;
[0039] Fig. 9 It is a block diagram of a terminal device shown in an exemplary embodiment of the present application.
[0040] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION
[0041] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0042] First, the terms involved in this application are explained:
[0043] Flink: An open source stream processing framework, the core of which is a distributed streaming data flow engine written in Java and Scala. Flink executes arbitrary streaming data programs in a data-parallel and pipelined manner. Flink's pipeline runtime system can execute batch and stream processing programs.
[0044] The application scenarios of the embodiments of the present application are explained below:
[0045] Figure 1FIG. 1 is an application scenario diagram of the data migration method provided in the embodiment of the present application. The data migration method provided in the embodiment of the present application can be applied to the scenario of data migration of heterogeneous databases. For example, Figure 1 As shown, the execution subject of the method provided in the embodiment of the present application can be a terminal device, and the terminal device communicates with the first storage device and the second storage device respectively through the network. Different types of databases are respectively running in the first storage device and the second storage device. For example, an Oracle database is running in the first storage device, and a Hive database is running in the second storage device. The databases in the first storage device and the second storage device are respectively used to store different business data. For example, the Oracle database in the first storage device stores product sales information, while the Hive database in the second storage device stores product parameter information. By executing the data migration method provided in the embodiment of the present application, the terminal device can migrate the data in the Oracle database of the first storage device to the Hive database in the second storage device, thereby realizing data migration of heterogeneous databases.
[0046] In the prior art, for database systems and tools such as hive, hbase, hdfs, and redis under the big data architecture, data is usually read, written, and managed based on one of them to achieve specific business needs. With the expansion of business and the continuous increase in data scale, it is necessary to merge and migrate data between different databases. However, due to the non-uniform data structures between different databases, when migrating data in different types of databases, it is necessary to first determine the mapping relationship between the two data structures, and then migrate the data of database A to database B based on this mapping relationship. However, due to the wide variety of databases, when it is necessary to migrate data of multiple types of databases, it is necessary to convert data for the mapping relationship between each source database and the target database, resulting in problems such as difficulty in data migration between different databases and low migration efficiency, which affects the data security and database management costs during data migration.
[0047] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0048] Figure 2 A flowchart of a data migration method provided by an embodiment of the present application is shown in FIG. Figure 2 As shown, the data migration method provided in this embodiment includes the following steps:
[0049] Step S101 : obtaining migration configuration information, the migration configuration information including configuration information of a first type of source database and configuration information of a second type of target database, the configuration information being used to perform read / write operations on data in databases of corresponding types.
[0050] Specifically, the source database is the database for migrating out data, and the target database is the database for migrating in data, wherein the type of the source database is different from the type of the target database, for example, the source database is an Hbase database, and the target database is a Redis database. Since the types of the source data and the target database are different, the data formats of the data stored in each database are also different, which results in that the data stored in the source database and the target database cannot be directly migrated. Furthermore, the migration configuration information is parameter information for implementing the data migration process between the source database and the target database, and the migration configuration information can be obtained by presetting the configuration data locally as the execution subject of the method of this embodiment. The migration configuration information includes the configuration information of the source database and the configuration information of the target database, wherein the configuration information of the source database is used to determine the type of the source data and read the source data from the source database; and the configuration information of the target database is used to determine the type of the target database, and write the processed source data to the target database, thereby completing the data migration process.
[0051] In a possible implementation, the configuration information in the migration configuration information includes a first information item and a second information item, wherein the first information item represents the database type of the corresponding database, and the second information item represents the database connection information of the corresponding database. The database type represented by the first information item includes database components under the big data architecture and traditional databases, such as Oracle, MySql, Hive, HDFS, Hbase, Redis, kafka, etc. The above database types are all commonly used database types that already exist in the prior art, and will not be introduced one by one here.
[0052] The second information item represents the database connection information of the database, including the connection address used to connect to and access the corresponding database, which can be described by a Json statement. For example, for an Oracle database, the database connection information is:
[0053] "jdbcUrl": ["jdbc:oracle:thin:@0.0.0.1:1521:oracle"]
[0054] The above database connection information can be used to access the Oracle database. For different databases, the database connection address is expressed differently. For example, for the Hbase database, the database connection information is:
[0055] "hbase.rootdir": "hdfs: / / ns1 / hbase"
[0056] The Hbase database can be accessed through the above database connection information. The access mentioned above may refer to reading or storing.
[0057] In addition, the implementation of the above database connection information has a specific grammatical structure and grammatical expression, so the above database connection information is only exemplary and may be different in different development environments, which will not be repeated here.
[0058] Step S102: based on the configuration information of the source database, source data is obtained from the source database, and the source data is saved as intermediate data, where the intermediate data is two-dimensional data based on Flink.
[0059] After obtaining the migration configuration information, firstly, the source database in the source database is obtained according to the configuration information of the source database in the migration configuration information. Specifically, according to the configuration information of the source database, the type of the source database is determined, and the corresponding connection address is obtained. Based on the connection address, the source database is accessed to read the source data in the database. Among them, exemplarily, the execution subject of the method provided in this embodiment is a terminal device. When the source database is set in the terminal device, the terminal device directly accesses the database to read the source data into the memory of the terminal device to generate intermediate data; when the source database is set outside the terminal device, such as an external storage device, the terminal device communicates with the external storage device through a communication network, and reads the source data in the source database located in the external storage device into the local memory of the terminal device to generate intermediate data.
[0060] Among them, further, the intermediate data has a specific data structure, that is, two-dimensional data based on Flink. The two-dimensional data refers to a data format that stores data in the form of key-value pairs, which can be applied to the Flink architecture, and can be processed based on the streaming or parallel data processing method of the Flink architecture. In the steps of this embodiment, in order to realize data migration between different types of databases, the source data in the source database is first converted into universal data in the form of two-dimensional data, that is, intermediate data. Then, the intermediate data in the form of two-dimensional data is converted into a data format matched by the target database, thereby realizing data format conversion and data migration between heterogeneous databases. At the same time, the streaming or parallel data processing capabilities of the Flink architecture are utilized to realize fast and continuous data migration, and the stability and security of the data transmission process are high, and it is not easy to cause data loss.
[0061] Step S103, based on the configuration information of the target database, convert the intermediate data into target data in the data format of the target database, and verify the target data to obtain a verification result. If the verification result is normal, store the target data in the target database.
[0062] Furthermore, after the intermediate data is generated, the relevant information of the target database is obtained based on the configuration information of the target database, specifically, for example, the type of the target database and the corresponding connection address of the target database, so that the terminal device can access the target database and perform a write operation. Afterwards, according to the data type of the target database, the intermediate data is converted into data that matches the data type of the target database, that is, the target data. That is, the data format of the target data matches the target database and can be written into the target database.
[0063] Afterwards, because the two conversion processes of "source data-intermediate data-target data" have been undergone in the previous steps, there may be conversion anomalies in the conversion process, which leads to inconsistency between the final generated target data and the source data, resulting in data loss and errors during the data migration process. Therefore, before writing the target data into the target database, the target data needs to be verified to ensure that no erroneous data is introduced during the above conversion process.
[0064] Specifically, in the process of streaming data reading and processing based on Flink, the source data includes multiple source data units, and the target data includes multiple target data units. Based on the Flink framework, each source data unit is stream-read, and each source data unit is converted twice to generate a corresponding target data unit. In this process, each time a source data unit is successfully read, the data identifier of the source data unit is recorded; each time a target data unit is successfully generated, the data identifier of the target data unit is recorded. After the conversion process is completed, the number of data identifiers of the source data units and the number of data identifiers of the target data units are counted, that is, the first number corresponding to the source data units and the second number of the target data units are obtained. Afterwards, the first quantity is compared with the second quantity. If the first quantity and the second quantity are consistent, it is considered that there is no error in the conversion process of "source data-intermediate data-target data", that is, the verification result is normal, and the target data can be written to the target database; conversely, if the first quantity and the second quantity are inconsistent, it is considered that there is an error in the conversion process of "source data-intermediate data-target data", that is, the verification result is abnormal, then the target data will not be written to the database, and the above steps can be re-executed to realize the re-migration of the source data in the source database, or a data migration failure prompt will be issued.
[0065] In this embodiment, by obtaining migration configuration information, the migration configuration information includes configuration information of a first type of source database and configuration information of a second type of target database, and the configuration information is used to perform read / write operations on data in databases of corresponding types; based on the configuration information of the source database, the source data is obtained from the source database, and the source data is saved as intermediate data, and the intermediate data is two-dimensional data based on Flink; based on the configuration information of the target database, the intermediate data is converted into target data in the data format of the target database, and the target data is verified to obtain a verification result. If the verification result is normal, the target data is stored in the target database.
[0066] Figure 3 A flowchart of a data migration method provided in another embodiment of the present application is shown in FIG. Figure 3 As shown, the data migration method provided in this embodiment is Figure 2 Based on the data migration method provided in the illustrated embodiment, steps S102-S103 are further refined, and the data migration method provided in this embodiment includes the following steps:
[0067] Step S201 : obtaining migration configuration information, the migration configuration information including configuration information of a first type of source database and configuration information of a second type of target database, the configuration information being used to perform read / write operations on data in databases of corresponding types.
[0068] Step S202, obtaining a preset data export object, the data export object includes export parameters, and the data export object is used to stream read each source data unit from the database indicated by the export parameters.
[0069] Step S203, based on the configuration information of the source database, obtain the export parameters of the data export object, call the data export object, and stream read each source data unit in the source database.
[0070] Specifically, an object refers to a function used to implement a specific function, and input parameters are usually set to adjust the output result and operation logic of the object. In this embodiment, the data export object is a function used to stream source data from a source database into memory. Among them, the data export object includes export parameters, which are used to indicate a specific database, that is, to control which database the data export object exports data from. The data export object is a function preset based on the Flink architecture, and its specific implementation method follows the grammatical rules of the Flink framework, which will not be described in detail here.
[0071] Furthermore, after obtaining the data export object, the data export object is called according to the configuration information of the source database as the export parameter, so as to achieve the purpose of the data export object streaming read the source data from the source database.
[0072] For example, Figure 4 As shown, the specific method of streaming reading each source data unit in the source database in step S203 includes looping the following steps until all source data units are read:
[0073] Step S2031, reading source data units in the source database through the data export object.
[0074] Step S2032: after reading the source data unit in the source database each time, detecting the reading result corresponding to the current source data unit, the reading result indicating whether the source data unit is read successfully.
[0075] Step S2033, if the current source data unit is read successfully, the data identifier corresponding to the current source data unit is recorded; if the source data unit is not read successfully, based on the verification identifier corresponding to the previous source data unit, the data reading object is re-called to read the current source data unit.
[0076] Exemplarily, after the source data units in the source database are read into the memory in a queue through the data export object, the corresponding intermediate data units in the memory are checked to see whether they are normal according to the preset check rules, for example, whether the corresponding intermediate data units are empty. If they are empty, it means that the source data units are not read successfully, that is, the reading result is normal. If they are not empty, it means that the source data units are read successfully, that is, the reading result is abnormal.
[0077] Furthermore, according to the reading result, when the reading result is normal, the data identifier corresponding to the current source data unit is recorded, wherein the data identifier is the unique identifier of the source data unit, and the source data unit can be located and identified according to the data identifier. When the reading result is abnormal, it means that an abnormality has occurred in the reading process. At this time, according to the verification identifier corresponding to the previous source data unit, the data reading object is re-called to read the current source data unit, thereby realizing breakpoint resume, ensuring that data is not lost during the data migration process, and improving the security of data migration.
[0078] In this embodiment, when it is determined that the reading result is normal, the data identifier corresponding to the current source data unit is recorded, with the purpose of recording the source data units that have been successfully read and identifying the source data units that have not been successfully read, thereby achieving breakpoint resumption of the source data units that have not been successfully read. At the same time, by marking the source data, data verification can be performed in subsequent steps, such as data marking of the source data, to further improve data security.
[0079] Step S204, writing the obtained source data units into the memory in sequence in the format of two-dimensional data to obtain intermediate data.
[0080] Furthermore, after reading the source data unit, the data export object writes the source data unit into the memory in two-dimensional data format based on the Flink architecture, and obtains the intermediate data stored in the memory for subsequent conversion to the data format corresponding to the target database. Figure 2 It is introduced in the illustrated embodiment and will not be described in detail here.
[0081] Step S205, obtaining a preset data import object, the data import object including import parameters, and the data import object is used to convert the intermediate data unit into target data matching the data format of the database indicated by the import parameters.
[0082] Step S206, based on the configuration information of the target database, obtain the import parameters of the data import object, call the data import object, and convert the intermediate data unit into the target data.
[0083] Specifically, the data import object is similar to the data export object, and is a function for writing data in memory to a file. Specifically, the data import object in this embodiment converts the intermediate data into target data in the target database format according to the target database indicated by the configuration information, and writes the target database. Its specific implementation principle is similar to that of the data export object, and will not be repeated here.
[0084] For example, Figure 5 As shown, the specific method of converting the intermediate data unit into the target data in step S206 includes looping the following steps until all the intermediate data units are converted into the target data:
[0085] Step S2061, converting the intermediate data units in the memory into corresponding target data units in a streaming manner through the data import object.
[0086] Step S2062: After each conversion of the intermediate data unit into the corresponding target data unit, a conversion result corresponding to the current target data unit is detected, where the conversion result indicates whether the current target data unit is successfully converted.
[0087] Step S2063, if the current target data unit is converted successfully, the data identifier corresponding to the current target data unit is recorded; if the current target data unit is not converted successfully, based on the verification identifier corresponding to the previous target data unit, the data import object is re-called to convert the intermediate data unit corresponding to the current target data unit.
[0088] Specifically, by calling the data import object, the intermediate data units maintained in the memory are stream-converted into corresponding target data units. After each conversion of the intermediate data units into corresponding target data units, the conversion result corresponding to the current target data unit is detected. If the conversion of the current target data unit is successful, the data identifier corresponding to the current target data unit is recorded; if unsuccessful, breakpoint conversion is performed based on the previous target data unit, and the current target data unit is converted again.
[0089] Step S207, obtaining a first quantity corresponding to the source data unit and a second quantity corresponding to the target data unit.
[0090] Step S208, when the first quantity is equal to the second quantity, obtain first sampling data and second sampling data, wherein the first sampling data includes multiple source data units, the second sampling data includes multiple target data units of the same number, and the data identifier of each source data unit corresponds to the same data identifier of each target data unit.
[0091] For example, after the conversion process of the target data is completed, the first number corresponding to the source data unit in the source database is compared with the second number of the target data unit. If they are inconsistent, it means that data loss has occurred during the migration process. Figure 2 The relevant steps in the illustrated embodiment have been described and will not be repeated here. However, when the first number corresponding to the source data unit and the second number of the target data unit are consistent, there is also a case of data error. Therefore, it is necessary to further verify the target data to ensure the correctness and security of the data. Specifically, when the first number is equal to the second number, multiple source data units are extracted from the source data as the first sampling data; and multiple corresponding target data units are extracted from the target data as the second sampling data. Among them, the source data unit in the first sampling data and the target data unit in the second sampling data are one-to-one corresponding, that is, the data identifier of the source data unit in the first sampling data and the data identifier of the target data unit in the second sampling data are one-to-one corresponding. The data identifier of the source data unit in the first sampling data and the data identifier of the target data unit in the second sampling data are the identification information recorded when the source data unit and the target data unit are successfully converted during the conversion process of the source data unit and the target data unit. Therefore, it can be ensured that the source data unit and the target data unit as the sampling data are corresponding and consistent.
[0092] More specifically, the steps of obtaining the first sampling data and the second sampling data include: randomly extracting N first data identifiers from the data identifiers corresponding to the source data units that have been successfully read; based on the N first data identifiers, determining N second data identifiers that correspond to or are the same as the N first data identifiers from the data identifiers corresponding to the target data units that have been successfully converted; and determining the target data units corresponding to the N second data identifiers as the second sampling data.
[0093] Step S209, performing verification based on the MD5 value of the first sampled data and the MD5 value of the second sampled data to obtain a verification result, wherein if the MD5 value of the first sampled data is the same as the MD5 value of the second sampled data, the verification result is normal.
[0094] Further, after obtaining the first sampled data and the corresponding second sampled data, the MD5 value of the first sampled data and the MD5 value of the second sampled data are calculated, and compared according to the calculation results to obtain a verification result. If the calculation results of the MD5 values are consistent, the verification result is normal. On the contrary, if the calculation results of the MD5 values are inconsistent, the verification result is abnormal. Among them, the calculation method of MD5 is a prior art known to those skilled in the art, and will not be described in detail here.
[0095] In this embodiment, the generated target data is verified through two steps based on the number of data identifiers and the MD5 value of the sampled data to ensure that the generated target data is identical to the source data, thereby ensuring data security and reliability during the data migration process.
[0096] Step S210, if the verification result is normal, the target data is stored in the target database; if the verification result is abnormal, abnormal information is output, and the abnormal information includes the first sample data and the second sample data with different MD5 values.
[0097] Furthermore, if the verification result is normal, the target data is stored in the target database to complete the data migration process; if the verification result is abnormal, an abnormal information is output to indicate that the data migration is abnormal, and the abnormal data is indicated, that is, the source data unit that causes the MD5 value in the first sampled data and the target data unit that causes the MD5 value to be different in the second sampled data. This allows the user to migrate the source data unit through other methods, such as manual conversion and copying, to ensure the integrity of the data migration.
[0098] Figure 6 A data verification process is provided in the embodiment of the present application, such as Figure 6As shown, after the source data in the source database is processed, two-dimensional data is generated, and then the target data is generated through the data import object. In the process of the data export object and the data import object processing the source data and the intermediate data, a first data identifier and a second data identifier representing the data processing record result are generated respectively. Verification is performed based on the first data identifier, the second data identifier and the corresponding sampling data to obtain the verification result. Then, according to the verification result, the target data is written into the target database or an abnormal information is issued.
[0099] In this embodiment, the implementation method of step S201 is the same as that of the present application. Figure 2 The implementation method of step S101 in the illustrated embodiment is the same and will not be described in detail here.
[0100] Figure 7 A schematic diagram of the structure of a data migration device provided in one embodiment of the present application is shown in FIG. Figure 7 As shown, the data migration device 3 provided in this embodiment includes:
[0101] A configuration module 31 is used to obtain migration configuration information, where the migration configuration information includes configuration information of a first type of source database and configuration information of a second type of target database, and the configuration information is used to perform read / write operations on data in databases of corresponding types;
[0102] A reading module 32 is used to obtain source data from a source database based on configuration information of the source database, and save the source data as intermediate data, where the intermediate data is two-dimensional data based on Flink;
[0103] The writing module 33 is used to convert the intermediate data into target data in the data format of the target database based on the configuration information of the target database, and verify the target data to obtain a verification result. If the verification result is normal, the target data is stored in the target database.
[0104] In one possible implementation, the source data includes multiple source data units, and the reading module 32 is specifically used to: obtain a preset data export object, the data export object includes export parameters, and the data export object is used to stream read each source data unit from the database indicated by the export parameters; based on the configuration information of the source database, obtain the export parameters of the data export object, and call the data export object to stream read each source data unit in the source database; write the obtained source data units into the memory in sequence in the format of two-dimensional data to obtain intermediate data.
[0105] In a possible implementation, the reading module 32 is further used to: detect a reading result corresponding to the current source data unit after each reading of a source data unit in the source database, the reading result indicating whether the source data unit has been read successfully; based on the reading result, if the current source data unit has been read successfully, record a data identifier corresponding to the current source data unit; if the source data unit has not been read successfully, re-call the data reading object to read the current source data unit based on the verification identifier corresponding to the previous source data unit.
[0106] In one possible implementation, the intermediate data includes multiple intermediate data units. When the writing module 33 converts the intermediate data into target data in the data format of the target database based on the configuration information of the target database, it is specifically used to: obtain a preset data import object, the data import object includes import parameters, and the data import object is used to convert the intermediate data unit into target data matching the data format of the database indicated by the import parameters; based on the configuration information of the target database, obtain the import parameters of the data import object, and call the data import object to convert the intermediate data unit into the target data.
[0107] In one possible implementation, the source data includes multiple source data units, and the target data includes multiple target data units. When verifying the target data, the writing module 33 is specifically used to: obtain a first quantity corresponding to the source data units, and a second quantity corresponding to the target data units; and obtain a verification result based on the first quantity and the second quantity, wherein if the first quantity is the same as the second quantity, the verification result is normal.
[0108] In a possible implementation, when the writing module 33 obtains the verification result based on the first quantity and the second quantity, it is specifically used to: when the first quantity is equal to the second quantity, obtain the first sampling data and the second sampling data, wherein the first sampling data includes multiple source data units, and the second sampling data includes multiple target data units of the same number, and the data identifier of each source data unit corresponds to the same data identifier of each target data unit; perform verification based on the MD5 value of the first sampling data and the MD5 value of the second sampling data to obtain the verification result, wherein if the MD5 value of the first sampling data is the same as the MD5 value of the second sampling data, the verification result is normal.
[0109] In a possible implementation manner, the configuration information includes a first information item and a second information item, wherein the first information item represents a database type of a corresponding database, and the second information item represents database connection information of the corresponding database.
[0110] The configuration module 31, the reading module 32 and the writing module 33 are connected in sequence. The data migration device 3 provided in this embodiment can execute the following steps: Figure 2-Figure 6The technical solutions of any of the method embodiments shown have similar implementation principles and technical effects, which will not be described in detail here.
[0111] Figure 8 A schematic diagram of an electronic device provided by an embodiment of the present application, such as Figure 8 As shown, the electronic device 4 provided in this embodiment includes: a processor 41, and a memory 42 communicatively connected to the processor 41.
[0112] Wherein, the memory 42 stores computer-executable instructions;
[0113] The processor 41 executes the computer execution instructions stored in the memory 42 to implement the present application. Figure 2-Figure 6 The data migration method provided in any one of the corresponding embodiments.
[0114] The memory 42 and the processor 41 are connected via a bus 43 .
[0115] For related instructions, please refer to Figure 2-Figure 6 The relevant descriptions and effects corresponding to the steps in the corresponding embodiments can be understood, and no further elaboration is made here.
[0116] An embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, and the computer program is executed by a processor to implement the present application. Figure 2-Figure 6 The data migration method provided in any one of the corresponding embodiments.
[0117] Among them, the computer readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, and the like.
[0118] One embodiment of the present application provides a computer program product, including a computer program, which implements the present application when executed by a processor. Figure 2-Figure 6 The data migration method provided in any one of the corresponding embodiments.
[0119] Fig. 9 It is a block diagram of a terminal device shown in an exemplary embodiment of the present application. The terminal device 800 can be a mobile phone, a computer, a digital broadcast terminal, a message transceiver device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0120] The terminal device 800 may include one or more of the following components: a processing component 802 , a memory 804 , a power component 806 , a multimedia component 808 , an audio component 810 , an input / output (I / O) interface 812 , a sensor component 814 , and a communication component 816 .
[0121] The processing component 802 generally controls the overall operation of the terminal device 800, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above-mentioned method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0122] The memory 804 is configured to store various types of data to support operations on the terminal device 800. Examples of such data include instructions for any application or method operating on the terminal device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0123] The power supply component 806 provides power to various components of the terminal device 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the terminal device 800.
[0124] The multimedia component 808 includes a screen that provides an output interface between the terminal device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the terminal device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.
[0125] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), and when the terminal device 800 is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 804 or sent via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.
[0126] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include but are not limited to: home button, volume button, start button, and lock button.
[0127] The sensor assembly 814 includes one or more sensors for providing various aspects of status assessment for the terminal device 800. For example, the sensor assembly 814 can detect the open / closed state of the terminal device 800, the relative positioning of the components, such as the display and keypad of the terminal device 800, and the sensor assembly 814 can also detect the position change of the terminal device 800 or a component of the terminal device 800, the presence or absence of contact between the user and the terminal device 800, the orientation or acceleration / deceleration of the terminal device 800, and the temperature change of the terminal device 800. The sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 may also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0128] The communication component 816 is configured to facilitate wired or wireless communication between the terminal device 800 and other devices. The terminal device 800 can access a wireless network based on a communication standard, such as WiFi, 3G, 4G, 5G or other standard communication networks, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0129] In an exemplary embodiment, the terminal device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-mentioned present application. Figure 2-Figure 6 The method provided in any one of the corresponding embodiments.
[0130] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, and the instructions can be executed by a processor 820 of a terminal device 800 to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0131] The embodiment of the present application also provides a non-temporary computer-readable storage medium, when the instructions in the storage medium are executed by the processor of the terminal device, the terminal device 800 can execute the above-mentioned present application Figure 2-Figure 6 The method provided in any one of the corresponding embodiments.
[0132] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of modules is only a logical function division. There may be other division methods in actual implementation, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.
[0133] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the application disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary techniques in the art that are not disclosed in the present application. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the following claims.
[0134] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
Claims
1. A data migration method, characterized in that: The method comprises: Acquire migration configuration information, the migration configuration information including configuration information of a first type of source database and configuration information of a second type of target database, the configuration information being used to perform read / write operations on data in databases of corresponding types; Based on the configuration information of the source database, source data is obtained from the source database, and the source data is saved as intermediate data, wherein the intermediate data is two-dimensional data based on Flink, and the two-dimensional data based on Flink is a data format applied to the Flink architecture and storing data in the form of key-value pairs, and the two-dimensional data based on Flink is general data; Based on the configuration information of the target database, convert the intermediate data into target data in the data format of the target database, verify the target data to obtain a verification result, and if the verification result is normal, store the target data in the target database; The source data includes a plurality of source data units, the target data includes a plurality of target data units, and verifying the target data includes: Obtaining a first quantity corresponding to the source data unit and a second quantity corresponding to the target data unit; When the first number is equal to the second number, obtaining first sampled data and second sampled data, wherein the first sampled data includes a plurality of source data units, the second sampled data includes a plurality of target data units of the same number, and a data identifier of each source data unit corresponds to and is the same as a data identifier of each target data unit; A verification is performed based on the MD5 value of the first sampled data and the MD5 value of the second sampled data to obtain a verification result, wherein if the MD5 value of the first sampled data is the same as the MD5 value of the second sampled data, the verification result is normal.
2. The method according to claim 1, characterized in that The source data includes a plurality of source data units. Based on the configuration information of the source database, the source data is obtained from the source database, and the source data is saved as intermediate data, including: Obtaining a preset data export object, the data export object including export parameters, the data export object being used to stream read each of the source data units from a database indicated by the export parameters; Based on the configuration information of the source database, the export parameters of the data export object are obtained, and the data export object is called to stream read each source data unit in the source database; The obtained source data units are sequentially written into the memory in the format of the two-dimensional data to obtain the intermediate data.
3. The method according to claim 2, characterized in that The method further comprises: After each reading of a source data unit in the source database, detecting a reading result corresponding to the current source data unit, the reading result indicating whether the source data unit is successfully read; According to the reading result, if the current source data unit is read successfully, recording the data identifier corresponding to the current source data unit; If the source data unit is not read successfully, the data reading object is re-called to read the current source data unit based on the verification identifier corresponding to the previous source data unit.
4. The method according to claim 1, characterized in that: The intermediate data includes a plurality of intermediate data units, and the converting of the intermediate data into target data in a data format of the target database based on the configuration information of the target database includes: Acquire a preset data import object, the data import object including import parameters, the data import object being used to convert the intermediate data unit into target data matching a data format of a database indicated by the import parameters; Based on the configuration information of the target database, the import parameters of the data import object are obtained, and the data import object is called to convert the intermediate data unit into the target data.
5. The method according to any one of claims 1 to 4, characterized in that: The configuration information includes a first information item and a second information item, wherein the first information item represents a database type of a corresponding database, and the second information item represents database connection information of the corresponding database.
6. A data migration device, characterized in that: include: A configuration module, used to obtain migration configuration information, wherein the migration configuration information includes configuration information of a first type of source database and configuration information of a second type of target database, wherein the configuration information is used to perform read / write operations on data in databases of corresponding types; A reading module, configured to obtain source data from the source database based on the configuration information of the source database, and save the source data as intermediate data, wherein the intermediate data is two-dimensional data based on Flink, and the two-dimensional data based on Flink is a data format applied to the Flink architecture and storing data in the form of key-value pairs, and the two-dimensional data based on Flink is general data; A writing module, configured to convert the intermediate data into target data in the data format of the target database based on the configuration information of the target database, and verify the target data to obtain a verification result, and if the verification result is normal, store the target data in the target database; The source data includes multiple source data units, and the target data includes multiple target data units. When the writing module verifies the target data, it is specifically used to: obtain a first quantity corresponding to the source data units, and a second quantity of the target data units; when the first quantity is equal to the second quantity, obtain first sampling data and second sampling data, wherein the first sampling data includes multiple source data units, and the second sampling data includes multiple target data units of the same quantity, and the data identifier of each source data unit corresponds to the same data identifier of each target data unit; verify according to the MD5 value of the first sampling data and the MD5 value of the second sampling data to obtain a verification result, wherein if the MD5 value of the first sampling data is the same as the MD5 value of the second sampling data, the verification result is normal.
7. An electronic device comprising: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the data migration method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the data migration method according to any one of claims 1 to 5 when executed by a processor.
9. A computer program product, characterized in that The invention comprises a computer program, which implements the data migration method according to any one of claims 1 to 5 when executed by a processor.
Citation Information
Patent Citations
Data processing method and device, computer readable storage medium and processor
CN113297326A