Data migration method and device, medium and program product
By executing write operations in the source database and recording write requests with the target key name in the key-value database, combined with asynchronous data processing and data cleansing, the problem of service interruption in MySQL database migration is solved, and efficient and flexible data migration and consistency management are achieved.
Patent Information
- Application Number
- CN202510666939.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-09-19
AI Technical Summary
The existing technology requires service interruption during the heterogeneous data migration process of MySQL databases and requires the version and configuration of the source and target ends to be consistent, which affects the user experience.
Data write operations are performed in the source database and write requests are recorded in the key-value database with the target key name. Target data is synchronized through asynchronous data processing. After the migration is complete, data is operated with the original key name. The target key name is scanned and cleaned to ensure data consistency.
It enables data migration without service interruption, improves migration efficiency, ensures data consistency and rollback, and allows users to flexibly choose to migrate data.
Smart Images

Figure CN120670401A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a data migration method, device, electronic device, computer-readable medium, and computer program product. Background Art
[0002] Based on existing solutions, database data migration, such as MySQL, typically involves two methods: physical migration and logical migration. Physical migration involves migrating data by making physical copies of database files, while logical migration involves performing data insert, update, and delete operations.
[0003] In heterogeneous data migration solutions for MySQL, the entire table is typically scanned starting with the primary key. After the structure is modified, the data is written to the new database. Once all data has been scanned, service is briefly interrupted to switch the data source to a heterogeneous database, such as a key-value store, and resume service. However, this approach prevents users from accessing the database while the MySQL service is down, and requires that the MySQL version and configuration remain consistent with the source during the downtime migration, impacting the user experience. Summary of the Invention
[0004] Multiple aspects of the present application provide a data migration method, apparatus, electronic device, computer-readable medium, and computer program product.
[0005] In one aspect of the present application, a data migration method is provided, wherein the method comprises:
[0006] In response to the data migration instruction, executing a first data processing strategy to process a data write request from the client, wherein the first data processing strategy performs a corresponding data write operation in the source database and records the data write request with a target key name in the key-value database;
[0007] Synchronize target data in the source data to a key-value database by performing asynchronous data processing, wherein the target data corresponds to a data segment selected by a user;
[0008] In response to an instruction to complete the data migration, switching the executed first data processing strategy to a second data processing strategy, wherein the second data processing strategy causes the data to be operated with the original key name in the key-value database;
[0009] By scanning the target key names in the key-value database, the target key names in the key-value database are cleaned.
[0010] In one aspect of the present application, a device for data migration is provided, wherein the device includes:
[0011] means for executing a first data processing strategy to process a data write request from a client in response to a data migration instruction, wherein the first data processing strategy performs a corresponding data write operation in a source database and records the data write request with a target key name in a key-value database;
[0012] means for synchronizing target data in the source data to a key-value database by performing asynchronous data processing, the target data corresponding to a data segment selected by a user;
[0013] means for switching, in response to an instruction to complete the data migration, from executing a first data processing strategy to a second data processing strategy, the second data processing strategy causing data to be manipulated with original key names in the key-value database;
[0014] A device for cleaning target key names in a key-value database by scanning the target key names in the key-value database.
[0015] Another aspect of the present application provides an electronic device, comprising:
[0016] at least one processor; and
[0017] a memory communicatively connected to the at least one processor; wherein,
[0018] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method of the embodiment of the application.
[0019] In another aspect of the present application, a computer-readable storage medium is provided, on which computer program instructions are stored. The computer program instructions can be executed by a processor to implement the method of the embodiment of the application.
[0020] In another aspect of the present application, a computer program product is provided, including a computer program, which implements the method of the embodiment of the present application when executed by a processor.
[0021] The solution provided by the embodiment of the present application performs corresponding data write operations in the source database and records the write request using the target key name in the key-value database during the migration process. The method of recording only the target key value reduces complex write operations, making data migration more convenient, simplifying the migration process, and improving migration efficiency. By adopting the data processing strategy of the embodiment of the present application to manage and synchronize data in the two database systems, data consistency and rollback are ensured, and smooth data migration and rollback are achieved without interrupting services. Data migration is performed based on user-selected data, so that users can flexibly choose which data to migrate, thereby improving data migration efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, a brief introduction will be given below to the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0023] Other features, objects and advantages of the present application will become more apparent upon reading the detailed description of non-limiting embodiments made with reference to the following drawings:
[0024] Figure 1 A schematic diagram of a data migration method according to an embodiment of the present invention is shown;
[0025] Figure 2 A schematic diagram showing an exemplary data migration process according to an embodiment of the present application is shown;
[0026] Figure 3 A schematic diagram of the structure of a device for data migration provided in an embodiment of the present application is shown;
[0027] Figure 4 A structural diagram of a device suitable for implementing the solution in the embodiments of the present application is shown.
[0028] The same or similar reference numerals in the drawings represent the same or similar components. DETAILED DESCRIPTION
[0029] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0030] In a typical configuration of the present application, the terminal and the equipment of the service network each include one or more processors (CPUs), input / output interfaces, network interfaces and memories.
[0031] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0032] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology for information storage. The information can be computer program instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc-read only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission medium that can be used to store information that can be accessed by a computing device.
[0033] Figure 1 A flow chart of a data migration method provided in an embodiment of the present application is shown, wherein the method comprises at least step S101, step S102, step S103 and step S104.
[0034] In practical scenarios, the execution subject of this method can be a network device or an application running on a network device. The network device includes, but is not limited to, a network host, a single network server, a set of multiple network servers, or a collection of computers based on cloud computing. Here, the cloud is composed of a large number of hosts or network servers based on cloud computing. Cloud computing is a type of distributed computing, a virtual computer composed of a group of loosely coupled computers.
[0035] In the process of migrating data from a source database to a key-value database, unlike traditional full-table scan migration or Binlog-based synchronization methods, the embodiment of the present application performs corresponding data write operations in the source database and uses the target key name to record the write request in the key-value database. During the migration process, the business does not need to be shut down, and the consistency and rollability of the data are ensured. Smooth data migration and rollback are achieved without interrupting the service, which is particularly suitable for high-concurrency scenarios.
[0036] For example, in the actual application scenario of video metadata update, it is necessary to update the video playback address, open status and other information. Since the query rate per second (QPS) requested by the system usually reaches the order of tens of thousands, it is necessary to continuously request to obtain the video playback address, as well as update the video address, whether the video is open and other information. During the whole process, the service cannot be suspended, otherwise it will affect the user experience. The embodiment of the present application can realize the extraction of data to form a record without affecting the original read and write operations, so as to facilitate the subsequent correction of the migrated data.
[0037] The source database and key-value database in the embodiment of the present application are heterogeneous data. In some embodiments, the source data is a relational database.
[0038] Reference Figure 1 In step S101, in response to a data migration instruction, a first data processing strategy is executed to process a data write request from a client.
[0039] The first data processing strategy executes a corresponding data writing operation in the source database and records the data writing request with a target key name in the key-value database.
[0040] The target key name is used to distinguish the old and new data versions during the migration process in the key-value database. In the embodiment of the present application, the target key name can be generated by a variety of methods, for example, adding a specific prefix or suffix to the original key name, inserting special characters in the middle, performing a hash operation on the original key name and using the resulting hash value as the corresponding target key name, etc.
[0041] According to one embodiment, the target key name is formed by adding a specific prefix to the original key name. For example, if the original key name is "key", the prefix "prefix_" is added to form "prefix_key", which is used as the target key name. Based on this method, the prefix can be used to quickly identify the key to be processed, facilitating subsequent data scanning.
[0042] The data write request includes inserting new data (INSERT), updating data (UPDATE), deleting data (DELETE), etc.
[0043] In a key-value database, a target key name is used to record a placeholder for a data operation, without storing specific values. The record may include key information such as the type of operation (e.g., insert, update, or delete) and the time of the operation, without focusing on the size or specific content of the actual stored value.
[0044] According to one embodiment, the first data processing strategy is implemented by a client agent. The method further includes step S105, and step S101 includes step S1011.
[0045] The client proxy is a middle-tier component deployed on the client side that handles database-related operations. For example, if the source database is MySQL, the client proxy intercepts SQL statements and key-value operations through the database connection pool or the ORM framework's hook mechanism to execute the first data processing strategy.
[0046] In step S105 , the agent program is deployed and configured in the client.
[0047] Specifically, a proxy program is deployed between the application and the database to ensure that it can intercept all database operation requests. According to the needs of the application, the proxy parameters are configured, such as database connection information, caching strategy, load balancing configuration, etc. It should be noted that the client directly connects to the database, that is, it is itself a backend service. After the upper layer requests the backend service, the service then accesses the database. The client proxy in the embodiment of the present application directly intercepts the request in the backend service.
[0048] In step S1011, the agent program is launched on the client to execute the first data processing strategy to process the data writing request of the client.
[0049] Among them, the data writing process of the client agent executing the first data processing strategy meets the following requirements: for each write request, the agent performs the following operations: when writing data to the source database, the data write operation is directly performed on the source database; when writing data to the key-value database, a specific prefix is added before the original key name of the data to be written to form a target key value, and the data write request is recorded with the target key name.
[0050] The data reading process in which the client agent executes the first data processing strategy includes: all data reading requests obtain data from the source database and the data reading result of the source database is used as the basis.
[0051] Optionally, before officially launching the client proxy, test the proxy to ensure that it can correctly intercept and handle requests and does not negatively impact the normal operation of the application. After confirming that the proxy is configured correctly and tested, officially launch the proxy to allow it to start handling the application's database requests.
[0052] Optionally, after the client agent is launched, the performance and stability of the agent are continuously monitored to optimize and adjust the agent program according to actual conditions.
[0053] Continue to refer to Figure 1 To illustrate, in step S102 , target data in the source data is synchronized to the key-value database by performing asynchronous data processing.
[0054] The target data corresponds to the data segment selected by the user.
[0055] The user queries the data in the source database and selects the target data segments to be migrated. The data in the source database can be queried using the database's own query tools (e.g., SQL queries) or third-party query tools. The method can select the target data segments manually, for example, manually selecting the data segments to be migrated through a graphical interface or command line tool. Alternatively, the method can select the target data segments automatically, for example, automatically selecting the corresponding data segments based on preset rules (e.g., data type, data volume, time range, etc.).
[0056] This method can be segmented to increase concurrency, significantly accelerating the data migration process. For example, in a MySQL scenario using a master-slave architecture, data can be segmented based on data IDs. Different data segments can be processed by different slave nodes. The same slave node can support multiple migration tasks concurrently, further increasing concurrency and accelerating data migration.
[0057] According to one embodiment, the method migrates target data in batches to a key-value database through asynchronous data processing based on a data range and migration times of the batch migration data set by a user.
[0058] Specifically, the method obtains the data range and migration frequency of the batch migration data set by the user. Specifically, the user can flexibly define the amount of data to be migrated each time (i.e., the data range) and the required migration batches (i.e., the number of migrations) based on actual needs and system resources. The user can determine the data range of the batch migration data based on the time range or data identifier. For example, the user can specify that the data volume for each migration is 1,000 records, and a total of 5 migrations are performed, thereby migrating a total of 5,000 target data records into the key-value database in batches.
[0059] Next, the method migrates the target data to the key-value database in batches, based on the user-defined data range and migration frequency. During each migration, the system extracts the appropriate amount of data based on the specified data range and transfers it to the key-value database via an asynchronous processing mechanism. This approach not only increases the flexibility of data migration but also optimizes the migration process based on actual user needs, ensuring efficient and reliable data migration.
[0060] In step S103, in response to the instruction to complete the data migration, the executed first data processing strategy is switched to a second data processing strategy, where the second data processing strategy enables the key-value database to operate data with the original key name.
[0061] After data migration is completed, in the key-value database, the key names of part of the data from the source database are the target key names, and the key names of the rest are the original key names.
[0062] The method may execute the operation of step S103 in response to an instruction to stop asynchronous data processing, or based on an operation instruction from a user, or at a preset time point.
[0063] It should be noted that after switching the first data processing strategy to the second data processing strategy, when inserting new data, the original key name is used for insertion; when performing data update (update) or deletion (delete) operations, if the original key name of the data does not exist in the target key name format in the key-value database, it means that the data has not been newly created or updated before migration, and the corresponding update operation or deletion operation can be performed directly; if the original key name of the data exists in the target key name format, it means that the data may have been newly created or updated before migration is performed. At this time, performing update operations or deletion operations may cause the data to be temporarily inaccurate. This problem can be fixed by performing subsequent scanning operations.
[0064] In step S104, the target key names in the key-value database are cleaned up by scanning the target key names in the key-value database.
[0065] Specifically, the key-value database is scanned based on the target key name, and the latest data is searched from the source database. If the data corresponding to the target key name does not exist in the source database, the relevant record of the target key name in the key-value database is deleted. If the data corresponding to the target key name exists in the source database, the latest value of the data in the source database is read and inserted into the key-value database with the original key name. After confirming that the data is correct, the relevant record of the target key name in the key-value database is deleted.
[0066] According to one embodiment, the method further includes step S106.
[0067] In step S106 , the consistency of the migrated data in the source database and the key-value database is verified. If there is any inconsistency, the difference is repaired based on the migrated data in the source database.
[0068] Specifically, the migration data in the source database and the key-value database are compared to determine whether the migration data in the source database and the key-value database are consistent. If they are inconsistent, the migration data in the key-value database is updated based on the data in the source database to ensure that the migration data in the source database and the key-value database are consistent.
[0069] Optionally, during the entire migration process, the migrated data is continuously verified for inconsistencies between the source database and the key-value database. If there are data inconsistencies, the differences are repaired based on the data in the source database to maintain data consistency.
[0070] According to one embodiment, if a problem occurs in the key-value database or data consistency needs to be verified, a rollback operation may be performed, and the method further includes step S107.
[0071] In step S107, in response to the data rollback instruction, a corresponding rollback operation is performed on the data that needs to be rolled back in the key-value database.
[0072] The rollback operation includes, but is not limited to, deleting data that has been migrated to the key-value database or overwriting data that needs to be rolled back with corresponding data in the source database.
[0073] Since the embodiment of the present application writes data changes to the source database and the key-value database in a dual-write manner during the data migration process, there is no need to perform additional backup operations on the source database when rolling back data.
[0074] According to the method of the embodiment of the present application, during the migration process, the corresponding data write operation is performed in the source database and the write request is recorded using the target key name in the key-value database. The method of only recording the target key value reduces complex write operations, making data migration more convenient, simplifying the migration process, and improving migration efficiency; by adopting the data processing strategy of the embodiment of the present application to manage and synchronize data in the two database systems, data consistency and rollback are ensured, and smooth data migration and rollback are achieved without interrupting service; data migration is performed based on user-selected data, so that users can flexibly choose which data to migrate, thereby improving data migration efficiency.
[0075] The following describes the data migration process of an embodiment of the present application with reference to an example.
[0076] Reference Figure 2 The example data migration scenario shown in the figure is MySQL.
[0077] Below Figure 2 Explain the names that appear in:
[0078] client: client;
[0079] proxy: agent;
[0080] async: asynchronous synchronization of data operations;
[0081] kv: key-value database (KV database);
[0082] mysql:MySQL database;
[0083] curd: represents the four basic actions of database operations. C stands for Create, which refers to the operation of inserting new data into the database; U stands for Update, which refers to the operation of modifying existing data in the database; R stands for Retrieve, which refers to the operation of reading data from the database; D stands for Delete, which refers to the operation of deleting data from the database.
[0084] The data migration process in this example consists of four stages. The following describes the process through steps P1 to P6, where step P1 corresponds to Figure 2 Steps P3 and P4 correspond to stage 3, and steps P5 and P6 correspond to stage 4.
[0085] Steps P1 to P6 are described below:
[0086] P1: Execute the data double-write strategy on the client;
[0087] The client deploys a proxy that implements a dual-write strategy, where all data changes are written simultaneously to MySQL and the key-value database.
[0088] The data writing process of the strategy includes: for each write request (such as INSERT, UPDATE, DELETE), the agent performs the following operations: Writing to MySQL: directly execute the original SQL statement to ensure that the main database data is updated in real time; Writing to the KV database: when writing data to the KV database, add the prefix "v2_" before the original key name "key" to form v2_key, so that the data write request is recorded with v2_key as the key value.
[0089] The data reading process of the strategy includes: all read requests first obtain data from MySQL, and the data reading results of MySQL are used as the basis.
[0090] P2: Synchronizes the target data in MySQL to the KV database through asynchronous data processing;
[0091] Select the start time for data migration and record the MySQL data location at that time. Based on the recorded MySQL data location, execute async to migrate the corresponding data to the KV database. For example, select time T1, record the corresponding data write location, and start migrating the data with id = 1000 from 1 to 1000 to the KV database.
[0092] P3: After the migration is complete, switch to the dual-write strategy so that the KV storage operates in normal key-value format;
[0093] After the data migration is complete, all data in MySQL will be stored in the KV database as v2_key and key. The client will switch to the dual-write strategy, allowing the KV database to perform data operations using the normal key format. This marks the completion of the main phase of the data migration, and the system will begin to use the KV database as the primary operation object.
[0094] After switching to the KV database and beginning normal key operations, update operations must be processed. If the updated key does not exist in the v2_key format in the KV database, the data was not updated before the migration, and the update or delete operation can be performed directly. However, if the updated key exists in the v2_key format, the data may have been created or updated before the migration. Performing an update or delete operation at this time may cause temporary data inaccuracy, which will be corrected by subsequent P4 operations.
[0095] P4: Clean up the v2_key in the KV storage;
[0096] To ensure data consistency, the v2_key in the KV database is scanned and the latest data is searched from MySQL. If the data corresponding to the v2_key does not exist in MySQL, the corresponding record in the KV database is deleted. If the data corresponding to the v2_key exists in MySQL, the latest value of the data in MySQL is read and inserted into the KV database using the key. After confirming the data is correct, the corresponding record is deleted.
[0097] P5: Verify data consistency and fix any discrepancies using MySQL as a benchmark.
[0098] During the entire migration process, the migrated data is continuously checked for inconsistencies in MySQL and KV. If there are inconsistencies, the differences are corrected based on the data in the MySQL database to maintain data consistency.
[0099] P6: Dual-write and dual-read data in MySQL and KV databases. Data can be switched between MySQL and KV databases or rolled back when necessary.
[0100] The method described in this example allows data to be migrated from MySQL to a KV database without service interruption, while ensuring data consistency and enabling rollback when necessary. For example, when applying the steps in this example to an e-commerce order data migration scenario, the e-commerce system client deploys an agent and writes order data. The agent writes to both MySQL and the KV database simultaneously, prefixing the original key names when writing to the KV database. Read requests prioritize MySQL data. During low-traffic periods, the target order data in MySQL is asynchronously migrated to the KV database in batches to minimize performance impact. After the data migration is complete and verified, the client switches to a dual-write strategy, and the KV database begins operating on data in normal key-value format. Prefixed key values in the KV store are scanned and compared with the MySQL data. After any inconsistencies are corrected, the prefixed key values are deleted. The two databases are continuously verified, with differences fixed using the MySQL data as the baseline. Dual-write and dual-read operations are maintained for a period of time, with the ability to switch back to MySQL or rollback data when necessary. By utilizing the dual-write strategy, asynchronous migration, and data verification and repair, consistent e-commerce order data migration can be achieved without downtime.
[0101] Figure 3 A schematic structural diagram of an apparatus for data migration provided in an embodiment of the present application is shown.
[0102] The device includes: a device for executing a first data processing strategy to process a client's data write request in response to a data migration instruction (hereinafter referred to as "first processing device 101"), a device for synchronizing target data in the source data to a key-value database by performing asynchronous data processing (hereinafter referred to as "asynchronous synchronization device 102"), a device for switching the executed first data processing strategy to a second data processing strategy in response to an instruction to complete data migration (hereinafter referred to as "second processing device 103"), and a device for cleaning target key names in the key-value database by scanning the target key names in the key-value database (hereinafter referred to as "record cleaning device 104").
[0103] Reference Figure 3 In response to the data migration instruction, the first processing device 101 executes the first data processing strategy to process the client's data write request.
[0104] The first data processing strategy executes a corresponding data writing operation in the source database and records the data writing request with a target key name in the key-value database.
[0105] The target key name is used to distinguish the old and new data versions during the migration process in the key-value database. In the embodiment of the present application, the target key name can be generated by a variety of methods, for example, adding a specific prefix or suffix to the original key name, inserting special characters in the middle, performing a hash operation on the original key name and using the resulting hash value as the corresponding target key name, etc.
[0106] According to one embodiment, the target key name is formed by adding a specific prefix to the original key name. For example, if the original key name is "key", the prefix "prefix_" is added to form "prefix_key", which is used as the target key name. Based on this method, the prefix can be used to quickly identify the key to be processed, facilitating subsequent data scanning.
[0107] The data write request includes inserting new data (INSERT), updating data (UPDATE), deleting data (DELETE), etc.
[0108] In a key-value database, a target key name is used to record a placeholder for a data operation, without storing specific values. The record may include key information such as the type of operation (e.g., insert, update, or delete) and the time of the operation, without focusing on the size or specific content of the actual stored value.
[0109] According to one embodiment, the first data processing strategy is implemented by a client agent. The apparatus further comprises an agent deployment device.
[0110] The client proxy is a middle-tier component deployed on the client side that handles database-related operations. For example, if the source database is MySQL, the client proxy intercepts SQL statements and key-value operations through the database connection pool or the ORM framework's hook mechanism to execute the first data processing strategy.
[0111] The agent deployment device deploys and configures the agent program in the client.
[0112] Specifically, a proxy program is deployed between the application and the database to ensure that it can intercept all database operation requests. According to the needs of the application, the proxy parameters are configured, such as database connection information, caching strategy, load balancing configuration, etc. It should be noted that the client directly connects to the database, that is, it is itself a backend service. After the upper layer requests the backend service, the service then accesses the database. The client proxy in the embodiment of the present application directly intercepts the request in the backend service.
[0113] The agent program is launched on the client, so that the first processing device 101 processes the client's data writing request by executing the first data processing strategy through the processing program.
[0114] Among them, the data writing process of the client agent executing the first data processing strategy meets the following requirements: for each write request, the agent performs the following operations: when writing data to the source database, the data write operation is directly performed on the source database; when writing data to the key-value database, a specific prefix is added before the original key name of the data to be written to form a target key value, and the data write request is recorded with the target key name.
[0115] The data reading process in which the client agent executes the first data processing strategy includes: all data reading requests obtain data from the source database and the data reading result of the source database is used as the basis.
[0116] Optionally, before officially launching the client proxy, test the proxy to ensure that it can correctly intercept and handle requests and does not negatively impact the normal operation of the application. After confirming that the proxy is configured correctly and tested, officially launch the proxy to allow it to start handling the application's database requests.
[0117] Optionally, after the client agent is launched, the performance and stability of the agent are continuously monitored to optimize and adjust the agent program according to actual conditions.
[0118] Continue to refer to Figure 3 To illustrate, in step S102 , the asynchronous synchronization device 102 synchronizes the target data in the source data to the key-value database by performing asynchronous data processing.
[0119] The target data corresponds to the data segment selected by the user.
[0120] The user queries the data in the source database and selects the target data segments to be migrated. The data in the source database can be queried using the database's own query tools (e.g., SQL queries) or third-party query tools. The method can select the target data segments manually, for example, manually selecting the data segments to be migrated through a graphical interface or command line tool. Alternatively, the method can select the target data segments automatically, for example, automatically selecting the corresponding data segments based on preset rules (e.g., data type, data volume, time range, etc.).
[0121] The device can be segmented to increase concurrency, significantly accelerating the data migration process. For example, in a MySQL scenario using a master-slave architecture, data can be segmented based on data IDs, with different data segments being processed by different slave nodes. The same slave node can support multiple migration tasks concurrently, further increasing concurrency and accelerating data migration.
[0122] According to one embodiment, the apparatus migrates target data in batches to the key-value database through asynchronous data processing based on a data range and migration times of the batch migration data set by a user.
[0123] Specifically, the device obtains the data range and migration frequency of the batch migration data set by the user. Specifically, the user can flexibly define the amount of data to be migrated each time (i.e., the data range) and the migration batches to be performed (i.e., the number of migrations) based on actual needs and system resource conditions. Among them, the user can determine the data range of the batch migration data based on the time range or data identifier. For example, the user can specify that the amount of data to be migrated each time is 1,000 records, and a total of 5 migrations are performed, thereby migrating a total of 5,000 target data into the key-value database in batches.
[0124] Next, the device migrates the target data to the key-value database in batches, based on the user-defined data range and migration frequency. During each migration, the system extracts the appropriate amount of data based on the specified data range and transfers it to the key-value database via an asynchronous processing mechanism. This approach not only increases the flexibility of data migration but also allows the migration process to be optimized based on the user's actual needs, ensuring efficient and reliable data migration.
[0125] Continue to refer to Figure 3 To illustrate, in response to the instruction to complete the data migration, the second processing means 103 switches the executed first data processing strategy to the second data processing strategy, where the second data processing strategy enables the key-value database to operate data with the original key name.
[0126] After data migration is completed, in the key-value database, the key names of part of the data from the source database are the target key names, and the key names of the rest are the original key names.
[0127] The operation of the second processing device 103 may be executed in response to an instruction to stop asynchronous data processing. Alternatively, the operation of the second processing device 103 may be executed based on an operation instruction from a user. Alternatively, the operation of the second processing device 103 may be executed when a preset time point is reached.
[0128] It should be noted that after switching the first data processing strategy to the second data processing strategy, when inserting new data, the original key name is used for insertion; when performing data update (update) or deletion (delete) operations, if the original key name of the data does not exist in the target key name format in the key-value database, it means that the data has not been newly created or updated before migration, and the corresponding update operation or deletion operation can be performed directly; if the original key name of the data exists in the target key name format, it means that the data may have been newly created or updated before migration is performed. At this time, performing update operations or deletion operations may cause the data to be temporarily inaccurate. This problem can be fixed by performing subsequent scanning operations.
[0129] The record cleaning device 104 cleans up the target key names in the key-value database by scanning the target key names in the key-value database.
[0130] Specifically, the key-value database is scanned based on the target key name, and the latest data is searched from the source database. If the data corresponding to the target key name does not exist in the source database, the relevant record of the target key name in the key-value database is deleted. If the data corresponding to the target key name exists in the source database, the latest value of the data in the source database is read and inserted into the key-value database with the original key name. After confirming that the data is correct, the relevant record of the target key name in the key-value database is deleted.
[0131] According to one embodiment, the method further comprises a consistency verification device.
[0132] The consistency verification device verifies the consistency of the migration data in the source database and the key-value database. If there is an inconsistency, the difference is repaired based on the migration data in the source database.
[0133] Specifically, the consistency verification device compares the migration data in the source database and the key-value database to determine whether the migration data is consistent in the source database and the key-value database. If inconsistent, the migration data in the key-value database is updated based on the data in the source database to ensure that the migration data in the source database and the key-value database are consistent.
[0134] Optionally, during the entire migration process, the consistency verification device continuously checks whether there are any inconsistencies in the migrated data between the source database and the key-value database. If there are data inconsistencies, the differences are repaired based on the data in the source database to maintain data consistency.
[0135] According to one embodiment, if a problem occurs in the key-value database or data consistency needs to be verified, a rollback operation can be performed, and the device further includes a rollback execution device.
[0136] In response to the data rollback instruction, the rollback execution device executes a corresponding rollback operation on the data that needs to be rolled back in the key-value database.
[0137] The rollback operation includes, but is not limited to, deleting data that has been migrated to the key-value database or overwriting data that needs to be rolled back with corresponding data in the source database.
[0138] Since the embodiment of the present application writes data changes to the source database and the key-value database in a dual-write manner during the data migration process, there is no need to perform additional backup operations on the source database when rolling back data.
[0139] According to the device of the embodiment of the present application, during the migration process, the corresponding data write operation is performed in the source database and the write request is recorded using the target key name in the key-value database. The method of recording only the target key value reduces complex write operations, making data migration more convenient, simplifying the migration process, and improving migration efficiency; by adopting the data processing strategy of the embodiment of the present application to manage and synchronize data in the two database systems, data consistency and rollback are ensured, and smooth data migration and rollback are achieved without interrupting services; data migration is performed based on user-selected data, so that users can flexibly choose which data to migrate, thereby improving data migration efficiency.
[0140] Based on the same inventive concept, an electronic device is also provided in an embodiment of the present application. The method corresponding to the electronic device may be the data migration method in the aforementioned embodiment, and its principle of solving the problem is similar to that of the method. The electronic device provided in an embodiment of the present application includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the methods and / or technical solutions of the aforementioned multiple embodiments of the present application.
[0141] The electronic device may be a user device, or a device formed by integrating a user device and a network device via a network, or an application running on the above device. The user device includes but is not limited to various terminal devices such as computers, mobile phones, tablets, smart watches, and bracelets. The network device includes but is not limited to network hosts, single network servers, multiple network server sets, or cloud computing-based computer collections, and can be used to implement some of the processing functions when setting an alarm. Here, the cloud is composed of a large number of hosts or network servers based on cloud computing (Cloud Computing), where cloud computing is a type of distributed computing, a virtual computer composed of a group of loosely coupled computers.
[0142] Figure 4The structure of a device suitable for implementing the method and / or technical solution in the embodiment of the present application is shown. The device 1200 includes a central processing unit (CPU) 1201, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1202 or the program loaded from the storage part 1208 into the random access memory (RAM) 1203. Various programs and data required for system operation are also stored in RAM 1203. CPU 1201, ROM 1202 and RAM 1203 are connected to each other through a bus 1204. Input / output (I / O) interface 1205 is also connected to bus 1204.
[0143] The following components are connected to the I / O interface 1205: an input section 1206 including a keyboard, a mouse, a touch screen, a microphone, an infrared sensor, and the like; an output section 1207 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), an LED display, an OLED display, and a speaker; a storage section 1208 including one or more computer-readable media such as a hard disk, an optical disk, a magnetic disk, and a semiconductor memory; and a communication section 1209 including a network interface card such as a LAN (Local Area Network) card, a modem, and the like. The communication section 1209 performs communication processing via a network such as the Internet.
[0144] In particular, the methods and / or embodiments of the present application can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for executing the method shown in the flowchart. When the computer program is executed by the central processing unit (CPU) 1201, the above-mentioned functions defined in the method of the present application are performed.
[0145] Another embodiment of the present application further provides a computer-readable storage medium having computer program instructions stored thereon, which can be executed by a processor to implement the methods and / or technical solutions of any one or more embodiments of the present application.
[0146] Specifically, the present embodiment can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination thereof. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device.
[0147] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0148] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0149] Computer program code for performing the operations of the present application can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0150] The flow chart or block diagram in the accompanying drawings illustrate the possible architecture, functions and operations of the equipment, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code include one or more executable instructions for realizing the logical function of the specification. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a dedicated system for hardware that performs the function or operation of the specification, or can be implemented with a combination of dedicated hardware and computer instructions.
[0151] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0152] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or page components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0153] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0154] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0155] The above-mentioned integrated unit implemented in the form of a software functional unit can be stored in a computer-readable storage medium. The above-mentioned software functional unit is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute some steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program code.
[0156] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
[0157] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in a device claim may also be implemented by a single unit or device through software or hardware. Terms such as "first" and "second" are used to indicate names and do not imply any particular order.
Claims
1. A data migration method, wherein: The method comprises: In response to the data migration instruction, executing a first data processing strategy to process a data write request from the client, wherein the first data processing strategy performs a corresponding data write operation in the source database and records the data write request with a target key name in the key-value database; Synchronize target data in the source data to a key-value database by performing asynchronous data processing, wherein the target data corresponds to a data segment selected by a user; In response to an instruction to complete the data migration, switching the executed first data processing strategy to a second data processing strategy, wherein the second data processing strategy causes the data to be operated with the original key name in the key-value database; By scanning the target key names in the key-value database, the target key names in the key-value database are cleaned.
2. The method according to claim 1, wherein The method further comprises: Verify the consistency of the migrated data in the source database and the key-value database. If any inconsistencies exist, fix the differences based on the migrated data in the source database.
3. The method according to claim 1, wherein The method implements the first data processing strategy through a client agent, and the method further includes: Deploy and configure the agent program in the client; The executing of the first data processing strategy to process the data write request of the client in response to the data migration instruction includes: The agent program is launched on the client to execute the first data processing strategy to process the data writing request of the client.
4. The method according to claim 3, wherein: The data writing process of the client agent executing the first data processing strategy meets the following requirements: When writing data to the source database, the data writing operation is performed directly on the source database; When writing data to a key-value database, a specific prefix is added to the original key name of the data to be written to form a target key value, and the data write request is recorded with the target key name.
5. The method according to claim 3, wherein The data reading process of the client agent executing the first data processing strategy meets the following requirements: All data read requests obtain data from the source database and are subject to the data read results of the source database.
6. The method according to claim 1, wherein The step of synchronizing the target data in the source data to the key-value database by performing asynchronous data processing includes: Based on the data range and migration frequency set by the user, the target data is migrated in batches to the key-value database through asynchronous data processing.
7. The method according to claim 1, wherein The cleaning of the target key name in the key-value database by scanning the target key name in the key-value database includes: Scan the key-value database based on the target key name and search for the latest data in the source database. If the source database finds that the data corresponding to the target key name does not exist, delete the relevant records of the target key name in the key-value database. If the source database contains data corresponding to the target key name, the latest value of the data in the source database is read and inserted into the key-value database with the original key name. After confirming that the data is correct, the relevant records of the target key name in the key-value database are deleted.
8. The method according to any one of claims 1 to 7, wherein The method further comprises: In response to the data rollback instruction, a corresponding rollback operation is performed on the data that needs to be rolled back in the key-value database.
9. A device for data migration, wherein: The device comprises: means for executing a first data processing strategy to process a data write request from a client in response to a data migration instruction, wherein the first data processing strategy performs a corresponding data write operation in a source database and records the data write request with a target key name in a key-value database; means for synchronizing target data in the source data to a key-value database by performing asynchronous data processing, the target data corresponding to a data segment selected by a user; means for switching, in response to an instruction to complete the data migration, from executing a first data processing strategy to a second data processing strategy, the second data processing strategy causing data to be manipulated with original key names in the key-value database; A device for cleaning target key names in a key-value database by scanning the target key names in the key-value database.
10. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.
11. A computer-readable medium having computer program instructions stored thereon, wherein the computer program instructions can be used by a processor to execute the method according to any one of claims 1 to 8.
12. A computer program product comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Cited By
Data table switching method, electronic equipment and program product
CN121579470A