Database migration method and device, equipment, medium and program product

By performing hash sharding and geographic coordinate partitioning on the database data, the issues of data accuracy and security during the migration of databases from old to new systems were resolved, achieving accuracy and stability in the database migration process.

CN120994642APending Publication Date: 2025-11-21INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511123644.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

During the switchover between old and new systems, the differences in data structure between the old and new system databases make it impossible to guarantee the accuracy and security of the migrated data.

Method used

By acquiring the data from the database to be migrated and sharding it into multiple shards according to matching partitioning rules, including hash sharding based on key-value information and partitioning based on the geographical coordinates and data characteristics of distributed database nodes, it is ensured that the data matches the architectural characteristics of the distributed database during the migration process.

Benefits of technology

It improves the accuracy and security of database migration, achieves uniform data distribution and load balancing, avoids single point of pressure, and ensures the stability and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120994642A_ABST
    Figure CN120994642A_ABST
Patent Text Reader

Abstract

The invention provides a database migration method which can be applied to the field of distributed technologies. The database migration method comprises the steps of obtaining first data of a to-be-migrated database; the first data is divided according to a first division rule, a plurality of first fragments are obtained, and the first division rule is matched with the architecture of the to-be-migrated database; dividing data included in each first fragment according to a second division rule to obtain a plurality of second fragments, the second division rule being matched with the architecture of the distributed database; and migrating the data included in each second fragment to a corresponding node of the distributed database. The invention further provides a database migration device and equipment, a medium and a program product.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of distributed technology, specifically to a database migration method, apparatus, device, medium, and program product. Background Technology

[0002] In some cases, as business grows rapidly, centralized systems can no longer meet business needs, necessitating a switch to a distributed system. During this switchover, differences in the data structures of the old and new system databases can compromise the accuracy and security of the migrated data. Summary of the Invention

[0003] In view of the above problems, this application provides database migration methods, apparatus, devices, media and program products.

[0004] According to a first aspect of this application, a database migration method is provided, comprising: obtaining first data of a database to be migrated; dividing the first data according to a first partitioning rule to obtain multiple first shards, wherein the first partitioning rule matches the architecture of the database to be migrated; dividing the data included in each first shard according to a second partitioning rule to obtain multiple second shards, wherein the second partitioning rule matches the architecture of a distributed database; and migrating the data included in each second shard to the corresponding node of the distributed database.

[0005] According to an embodiment of this application, the first data is divided according to a first partitioning rule to obtain multiple first fragments, including: performing hash partitioning on the first data according to the key-value information of the first data to obtain multiple first fragments.

[0006] According to an embodiment of this application, the data included in each first partition is divided according to a second partitioning rule to obtain multiple second partitions, including: dividing the data included in each first partition according to the geographical coordinates of each node of the distributed database and the data characteristics of the first data to obtain multiple second partitions.

[0007] According to an embodiment of this application, the data included in each first partition is divided into multiple second partitions based on the geographical coordinates of each node in the distributed database and the data characteristics of the first data. This includes: generating data identifiers for the data included in each first partition based on the geographical coordinates of each node in the distributed database and the data characteristics of the first data; and dividing the data included in each first partition into multiple second partitions based on the data identifiers.

[0008] According to an embodiment of this application, the method further includes: obtaining second data, wherein the second data is data generated by the database to be migrated during the first data migration process; dividing the second data according to a first partitioning rule to obtain multiple third shards; dividing the data included in each third shard according to a second partitioning rule to obtain multiple fourth shards; and migrating the data included in each fourth shard to the corresponding node of the distributed database.

[0009] According to an embodiment of this application, obtaining the second data includes: obtaining the second data based on the relevant table operations of the database to be migrated during the first data migration process.

[0010] According to an embodiment of this application, the method further includes: in response to the migration of the first data to the distributed database, after waiting for a preset time, controlling the migration of the second data to the distributed database.

[0011] A second aspect of this application provides a database migration apparatus, comprising: an acquisition module for acquiring first data of a database to be migrated; a first partitioning module for partitioning the first data according to a first partitioning rule to obtain multiple first shards, wherein the first partitioning rule matches the architecture of the database to be migrated; a second partitioning module for partitioning the data included in each first shard according to a second partitioning rule to obtain multiple second shards, wherein the second partitioning rule matches the architecture of a distributed database; and a migration module for migrating the data included in each second shard to the corresponding node of the distributed database.

[0012] A third aspect of this application provides an electronic device comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method described above.

[0013] A fourth aspect of this application also provides a computer-readable storage medium having a computer program or instructions stored thereon, which, when executed by a processor, implement the steps of the above-described method.

[0014] The fifth aspect of this application also provides a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described method. Attached Figure Description

[0015] The above-mentioned contents, other objects, features and advantages of this application will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:

[0016] Figure 1 The illustrations depict application scenarios of database migration methods, apparatuses, devices, media, and program products according to embodiments of this application.

[0017] Figure 2 A flowchart illustrating a database migration method according to an embodiment of this application is shown schematically;

[0018] Figure 3 A flowchart illustrating a database migration method according to another embodiment of this application is shown schematically;

[0019] Figure 4 A schematic diagram illustrating the structure of a database migration apparatus according to an embodiment of this application is shown; and

[0020] Figure 5 A block diagram schematically illustrates an electronic device suitable for implementing a database migration method according to an embodiment of this application. Detailed Implementation

[0021] The embodiments of this application will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of this application. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of this application for ease of explanation. However, it will be apparent that one or more embodiments may be implemented without these specific details. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of this application.

[0022] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The terms “comprising,” “including,” etc., as used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0023] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0024] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).

[0025] In the technical solution of this application, the user information (including but not limited to user personal information, user image information, user device information, such as location information) and data (including but not limited to data used for analysis, stored data, and displayed data) involved are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of related data all comply with relevant laws, regulations, and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entry points for users to choose to authorize or refuse.

[0026] In some examples, as banking services continue to develop, internet scenarios and payment services continue to grow rapidly, and the volume of quick payment transactions has repeatedly reached new highs, the original traditional centralized system can no longer meet the needs of the business, so it is necessary to switch the centralized system to a distributed system.

[0027] During the switchover between old and new systems, it is necessary to consider the data synchronization and migration between the two systems. Due to potential differences in the data structures of the old and new systems, and the time lag between the switchover and replacement, the accuracy and security of the migrated data cannot be guaranteed during the database migration process.

[0028] In view of this, embodiments of this application provide a database migration method.

[0029] Figure 1 The illustration shows an application scenario diagram of a database migration method, apparatus, device, medium, and program product according to embodiments of this application.

[0030] like Figure 1 As shown, application scenario 100 according to this embodiment may include the financial technology field. Network 104 is used as a medium to provide a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables, etc.

[0031] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).

[0032] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0033] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0034] It should be noted that the database migration method provided in this application embodiment can generally be executed by server 105. Correspondingly, the database migration device provided in this application embodiment can generally be located in server 105. The database migration method provided in this application embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the database migration device provided in this application embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105.

[0035] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0036] The following will be based on Figure 1 The described scene, through Figures 2-5 The database migration method according to the embodiments of this application will be described in detail.

[0037] Figure 2 A flowchart illustrating a database migration method according to an embodiment of this application is shown schematically.

[0038] like Figure 2 As shown, the database migration method of this embodiment includes operations S210 to S240, and the database migration method can be executed by a server.

[0039] In operation S210, the first data of the database to be migrated is obtained.

[0040] In embodiments of this application, the database to be migrated is a database that requires data migration. The first data may be at least a portion of the data in the database to be migrated.

[0041] In the embodiments of this application, the first data may be existing data in the database to be migrated. Existing data refers to the data held by the database to be migrated at the target time, and the existing data may be all the data included in the database to be migrated at the target time. Here, the target time refers to the time when the database migration method of this embodiment begins to be executed, that is, the target time may refer to the time when the migration of the first data of the database to be migrated begins.

[0042] In operation S220, the first data is divided according to the first partitioning rule to obtain multiple first shards, wherein the first partitioning rule matches the architecture of the database to be migrated.

[0043] In the embodiments of this application, the first partitioning rule is matched with the architecture of the database to be migrated. The first partitioning rule is a preset sharding rule determined based on the characteristics of the data stored in the database to be migrated.

[0044] For example, the first partitioning rule may include partitioning the first data based on the customer ID. The first partitioning rule may include partitioning the first data based on the transaction card number corresponding to the first data. The first partitioning rule may include partitioning the first data based on the key-value information of the first data.

[0045] In the embodiments of this application, sharding is a data distribution technique used to split a large database into multiple smaller, independent sub-databases (called shards), each shard storing a portion of the data. For example, the number of the first shards could be 16.

[0046] In embodiments of this application, the number of first shards can be determined based on the amount of data in the first data. For example, when the amount of data in the first data is large, it indicates that the system is under greater pressure, and the number of first shards can be increased.

[0047] Through the above operations, centralized data can be transformed into a distributed, user-friendly form. Essentially, it is a preprocessing layer for architecture matching, which not only continues the architecture rules of the database to be migrated, but also directly responds to the scalability requirements of the distributed database through a sharding mechanism.

[0048] In the embodiments of this application, multiple first shards can be stored in a primary queue, enabling the first data of the database to be migrated to be uniformly and discretely saved to the primary message queue. Through this stage of processing, the transformation and migration of the first data of the database to be migrated to distributed messages can be achieved. Simultaneously, the old system (the centralized system corresponding to the database to be migrated) and the new system (the distributed system corresponding to the distributed database) can be decoupled. The primary queue acts as a buffer layer, allowing for inconsistent processing speeds between the old and new systems. The old system can quickly write changes to the primary queue, while the new system can read and process data from the queue according to its own consumption capacity.

[0049] In operation S230, the data included in each first partition is divided according to the second partitioning rule to obtain multiple second partitions, wherein the second partitioning rule matches the architecture of the distributed database.

[0050] In the embodiments of this application, the second partitioning rule matches the architecture of the distributed database. The second partitioning rule is a preset sharding rule determined based on the characteristics of the distributed database.

[0051] For example, the second partitioning rule could include partitioning the data in each first shard based on the geographical location of each node in the distributed database. The second partitioning rule could also include partitioning the data in each first shard based on the table type in the distributed database.

[0052] In embodiments of this application, the data included in each first partition is divided according to a second partitioning rule to obtain multiple second partitions. The number of second partitions can be the same as the number of nodes included in the distributed database, and each second partition corresponds to a node included in the distributed database. Dividing the data included in each first partition according to the second partitioning rule means that at least a portion of the first data included in each first partition is partitioned, and the first data is assigned to the corresponding second partition.

[0053] In operation S240, the data included in each second shard is migrated to the corresponding node of the distributed database.

[0054] In the embodiments of this application, each second shard corresponds to a node in the distributed database. Migrating the data included in each second shard to the corresponding node in the distributed database allows the data included in the second shard to be migrated to the corresponding node, thereby enabling the migrated data to match the architectural characteristics of the new distributed system.

[0055] Through the embodiments of this application, dividing the first data according to the first partitioning rule can achieve preliminary load balancing. Distributing the data evenly through sharding avoids single-point pressure and improves processing capacity. Multiple first shards act as a buffer layer, allowing for differences in processing speed between the old system corresponding to the database to be migrated and the new system corresponding to the distributed database. Dividing the data included in each first shard according to the second partitioning rule yields multiple second shards. These second shards form the core transformation layer for upgrading data from physical sharding to business-oriented distribution, ensuring that the migrated data perfectly matches the architectural characteristics of the new distributed system. The database migration method of this embodiment can effectively improve the accuracy and security of migrated data.

[0056] In some embodiments, dividing the first data according to a first partitioning rule to obtain multiple first fragments includes: performing hash partitioning on the first data according to the key-value information of the first data to obtain multiple first fragments.

[0057] In the embodiments of this application, key-value information is a special data item used to identify records in a database table. It typically consists of one or more fields and is used to ensure the uniqueness, integrity, and query efficiency of the data. The key-value information of the first data can be selected from fields with high business relevance.

[0058] For example, the key-value information of the first data could be the customer ID corresponding to the first data. The key-value information of the first data could also be the transaction card number corresponding to the first data.

[0059] In the embodiments of this application, hash sharding refers to the uniform distribution of data to different shards using a hash function. The first data is hash-sharded based on its key-value information to obtain multiple first shards; that is, the hash value of the first data is calculated using a preset hash function, and the first data is allocated to the corresponding first shard based on its hash value.

[0060] For example, the first shard has 16 shards. Based on the customer ID of the first data, the hash value of the first data is calculated using a preset hash function, and the corresponding shard number is determined by the remainder obtained by dividing the calculated hash value by 16.

[0061] Through the embodiments of this application, hash sharding of the first data based on the key-value information of the first data can achieve uniform distribution of the first data and avoid single-point bottlenecks. First data with the same unique key-value will be assigned to the same shard, thereby ensuring the sequentiality of these operations. At the same time, multiple first shards serve as intermediate message queues, decoupling data capture and subsequent processing, and improving the scalability and fault tolerance of the database migration method of this embodiment.

[0062] In some embodiments, the data included in each first partition is divided according to a second partitioning rule to obtain multiple second partitions, including: dividing the data included in each first partition according to the geographical coordinates of each node of the distributed database and the data characteristics of the first data to obtain multiple second partitions.

[0063] In embodiments of this application, the data included in each first shard can be divided according to the geographical coordinates of each node in the distributed database. For example, logical partitions (such as North China) can be divided according to the geographical coordinates of each node in the distributed database. Each node in the distributed database corresponds to a second shard.

[0064] For example, a distributed database may include nodes located in North China. Data from North China within each first shard can be partitioned into the second shard corresponding to that node.

[0065] In the embodiments of this application, the logical partition of the data included in the first shard can be calculated from the dimension of routing set characteristics based on the region / business line hash. Then, the data can be divided into corresponding nodes of the distributed database according to the logical partition of the data. The logical partition of the corresponding node of the distributed database is the same as the logical partition of the data that was divided in the past.

[0066] Here, the routing set characteristics can be the path information of the network address corresponding to the data. For example, data can be uploaded from the Northeast region to the database to be migrated. Data can also be uploaded from the East China region to the database to be migrated. Since the nodes of the distributed database are located in different geographical locations, the data in the East China region can be distributed to nodes located in East China.

[0067] In embodiments of this application, the data included in each first shard can be divided according to the data characteristics of the first data. The data characteristics of the first data may include table dimension characteristics of the data. For example, the data included in each first shard can be vertically split according to business, and the data can be divided into corresponding business tables (such as payment tables / account tables). Table dimension characteristics are used to classify the data according to business tables, and different tables are divided into different second shards.

[0068] In embodiments of this application, the data characteristics of the first data may include unique key characteristics. For example, the unique key characteristic may be the transaction identification (ID) characteristic corresponding to the data. By hashing the transaction ID corresponding to the data, the specific partition of the data in the second shard can be determined. The unique key characteristic can also be obtained by performing a secondary hash on the key-value information of the data included in each first shard.

[0069] In the embodiments of this application, the geographical coordinates of each node of the distributed database and the data characteristics of the first data can be combined to determine the second shard corresponding to the data included in each first shard, and the data included in each first shard can be divided and the data included in the first shard can be divided into the corresponding second shard.

[0070] For example, table dimension features, routing set features, and unique key features can be combined to divide the data included in each first shard, forming multiple second shards with multiple dimensions.

[0071] Through the embodiments of this application, the operation of this embodiment can upgrade data from physical shards to a core transformation layer of business-oriented distribution, so that the migrated data can match the architectural characteristics of the new distributed system, thereby improving the reliability of the database migration method of this embodiment.

[0072] In some embodiments, the data included in each first partition is divided into multiple second partitions based on the geographical coordinates of each node in the distributed database and the data characteristics of the first data. This includes: generating data identifiers for the data included in each first partition based on the geographical coordinates of each node in the distributed database and the data characteristics of the first data; and dividing the data included in each first partition into multiple second partitions based on the data identifiers.

[0073] In the embodiments of this application, data identifiers for the data included in each first shard are generated based on the geographical coordinates of each node in the distributed database and the data characteristics of the first data. For example, multidimensional coordinates can be generated using three elements: table dimension features, routing set features, and unique key features. The table dimension represents a pre-defined table type mapping rule used to determine which business table (e.g., payment table, account table) the data belongs to. The routing set feature represents the hash result of the business attribute (region / cluster), and the routing set feature can be a logical partition divided according to business rules (such as region, business line). The unique key feature represents the key-value secondary hash used to determine the final partition, such as a transaction ID, for further subdivision.

[0074] In the embodiments of this application, a multi-dimensional queue identifier can be generated based on a combination of three elements, such as "table + route set + unique key partition". Then, the data included in each first partition can be divided according to the data identifier to obtain multiple second partitions.

[0075] For example, data with the same queue identifier can be written to the same second shard.

[0076] Through the embodiments of this application, the data included in each first fragment is divided into multiple second fragments according to the data identifier, which can improve the accuracy of data division and thus improve the accuracy and reliability of the database migration method of this embodiment.

[0077] Figure 3 A flowchart illustrating a database migration method according to another embodiment of this application is shown.

[0078] like Figure 3 As shown, the database migration method of this embodiment includes operations S310 to S340, and the database migration method can be executed by a server.

[0079] In operation S310, the second data is obtained, wherein the second data is the data generated by the database to be migrated during the migration of the first data.

[0080] In operation S320, the second data is divided according to the first partitioning rule to obtain multiple third fragments.

[0081] In operation S330, the data included in each third partition is divided according to the second partitioning rule to obtain multiple fourth partitions.

[0082] In operation S340, the data contained in each fourth shard is migrated to the corresponding node of the distributed database.

[0083] In embodiments of this application, the second data is data generated by the database to be migrated during the first data migration process. The second data may be incremental data of the database to be migrated during the first data migration process.

[0084] In some examples, during the process of migrating data from the database to be migrated to the distributed database, the old system corresponding to the database to be migrated generates incremental data (secondary data). While maintaining the normal operation of the old system and the new system (the system corresponding to the distributed database), it is impossible to safely and accurately migrate the incremental data from the database to be migrated to the distributed database.

[0085] In the embodiments of this application, the second data (incremental data) and the first data (existing data) are migrated separately. The migration step of the second data is not started until the first data is migrated to the distributed database and verification is completed, which can ensure the production security of the system's external services. After the first data migration is completed and verification is completed, the second data is migrated to the distributed database through incremental supplementation, which can ensure that the migration of the second data does not affect the operation of the system.

[0086] In the embodiments of this application, the first partitioning rule is matched with the architecture of the database to be migrated. The first partitioning rule is a preset sharding rule determined based on the characteristics of the data stored in the database to be migrated.

[0087] In embodiments of this application, the number of third shards can be determined based on the amount of data in the second data. For example, when the amount of data in the second data is large, it indicates that the system is under greater pressure, and the number of third shards can be increased.

[0088] In the embodiments of this application, the process of dividing the second data into multiple third fragments according to the first partitioning rule is similar to operation S220, and will not be described in detail here.

[0089] In the embodiments of this application, the second partitioning rule matches the architecture of the distributed database. The second partitioning rule is a preset sharding rule determined based on the characteristics of the distributed database.

[0090] In the embodiments of this application, the process of dividing the data included in each third partition into multiple fourth partitions according to the second partitioning rule is similar to operation S220, and will not be described in detail here.

[0091] In embodiments of this application, the number of fourth shards can be the same as the number of nodes included in the distributed database, with each fourth shard corresponding to a node in the distributed database. The data included in each third shard is divided according to the second partitioning rule; that is, the second data included in each third shard is partitioned, and the second data is assigned to the corresponding fourth shard.

[0092] In the embodiments of this application, each fourth shard corresponds to a node in the distributed database. Migrating the data included in each fourth shard to the corresponding node in the distributed database allows the data included in the fourth shard to be migrated to the corresponding node, thereby enabling the second data to match the architectural characteristics of the new distributed system.

[0093] By using the same sharding rules to divide the second data through the embodiments of this application, the continuity of the second data can be guaranteed, enabling the second data to be allocated to the corresponding nodes of the distributed database, thereby improving the reliability and stability of the database migration method of this embodiment.

[0094] In some embodiments, obtaining the second data includes: obtaining the second data based on the relevant table operations of the database to be migrated during the first data migration process.

[0095] In the embodiments of this application, the related table operations of the database to be migrated refer to the basic data management operations performed on the table structure of the database to be migrated, which may include types such as adding, deleting, modifying, and querying. These operations are the basis for the management and data maintenance of the database to be migrated.

[0096] In the embodiments of this application, the relevant table operations of the database to be migrated during the first data migration process are obtained, that is, the data operation records of the database to be migrated during the first data migration process are obtained.

[0097] For example, related table operations may include data manipulation table information, which represents the database table name that has been changed. For example, data manipulation table information may include information from the customer account table.

[0098] For example, related table operations may include data operation type information, which represents the category of database operation. For example, data operation type information may include "update account balance".

[0099] For example, related table operations may include data change field information, which represents the specific field that has been modified and its old and new values. For example, data change field information may include information about changes to account balances.

[0100] For example, related table operations may include additional information. This additional information may include metadata such as timestamps and transaction IDs.

[0101] In the embodiments of this application, after obtaining the relevant table operations of the database to be migrated during the first data migration process, a log parser can be used to read the log files corresponding to the relevant table operations, parse the binary data into structured records, and obtain a structured file. A data filtering layer is then used to filter the tables / operation types / fields in the structured file according to preset rules to obtain the filtering results. Finally, a text converter is used to convert the filtering results into text information to obtain the second data.

[0102] By using the embodiments of this application, the second data is obtained based on the relevant table operations of the database to be migrated during the first data migration process, which can improve the accuracy of the obtained second data and thus improve the reliability of the database migration method of this embodiment.

[0103] In some embodiments, the method further includes: in response to the migration of the first data to the distributed database, controlling the migration of the second data to the distributed database after waiting for a preset time.

[0104] In some examples, after migrating the first data from the database to be migrated to the distributed database, the new system corresponding to the distributed database needs preparation time such as initialization and loading of routing rules. During the migration of the first data, the second data will continue to accumulate in the database to be migrated. Since the distributed database is still in the initialization phase or the preparation phase such as loading routing rules, if the database to be migrated starts the migration step of the second data immediately after obtaining the second data, it will affect the normal operation of the distributed database.

[0105] In embodiments of this application, in response to the migration of first data to the distributed database, the migration of second data to the distributed database is controlled after a preset time has elapsed. The preset time can be determined based on the historical average migration duration.

[0106] In the embodiments of this application, the first predicted time of the migration of the first data can be calculated by the historical duration of the historical migration data and the data volume of the first data being migrated at present, and the initialization time of the distributed database can be predicted. The first predicted time is added to the initialization time to obtain the preset time.

[0107] For example, if the predicted migration time for the first data is 2 hours based on historical data and the initialization time for the distributed database is 10 minutes, the preset time can be set to 2 hours and 10 minutes.

[0108] In the embodiments of this application, the migration progress of the first data (such as the number of remaining tables) can be monitored in real time, and the preset time can be adjusted. If the migration times out, the cycle is automatically extended (e.g., from 2 hours to 3 hours).

[0109] In embodiments of this application, when the backlog of the second data in the database to be migrated exceeds a preset threshold, the number of the first shards can be expanded. For example, when the backlog of the second data reaches 80% of the threshold, expansion is triggered (e.g., the number of shards is adjusted from 16 to 32).

[0110] In the embodiments of this application, the migration progress of the second data can be monitored in real time. If there is a delay in the data migration process of the second data, an alarm message can be generated and sent to the user terminal.

[0111] By setting a relatively safe second data retention period through the embodiments of this application, the tidal impact problem in financial-grade data migration can be effectively solved, the impact of the second data migration on the system can be avoided, and the reliability and stability of the database migration method of this embodiment can be improved.

[0112] In some embodiments, the database migration method may further include the following steps.

[0113] (1) Old system data (second data) capture stage. Based on the database change log of the old system (the system corresponding to the database to be migrated, such as a centralized system), the real-time data replication system automatically collects the relevant table operations of the old system database and saves the changed content in text form to the database to be migrated, thus obtaining the second data.

[0114] (2) Primary Queue Discrete Stage. Based on the preset number of message fragments (based on data volume and system pressure assessment, e.g., 16 fragments), the second data is fragmented and stored according to the key-value information of the data, so as to achieve uniform and discrete storage of the second data in the primary message queue. Through this stage of processing, the second data is transformed and migrated into distributed messages (primary message queue).

[0115] (3) Data Routing and Conversion Stage. During the second data migration process, the second data distribution can be hashed according to the fields specified by the business (such as customer number or transaction card number), and stored in shards according to the hash value. To ensure the continuity of subsequent incremental data, the second data must be consistent with the sharding rules of the first data. Therefore, the second data cannot be entered into the database in the same way as the first data migration. A data routing and conversion mechanism needs to be introduced to calculate the storage location of the second data in real time according to the sharding calculation rules of the first data migration (such as hashing by customer number) to realize the redistribution of the second data.

[0116] (4) Second-level queue hashing stage. Based on the results of data routing transformation, the application routing layer performs secondary discretization on the consumed data records. Through elements such as tables, routing sets, and unique record keys, a multi-dimensional second-level message queue is formed to achieve standardized, unified, uniform, and efficient distribution of consumer cluster data.

[0117] (5) New system data loading stage. Based on the technical architecture of the distributed system, the new system (the system corresponding to the distributed database, such as the distributed system) realizes the rapid consumption of secondary queue messages through concurrent consumption threads of different tables, different shards, and different databases, and realizes near real-time synchronization of the old system change records to the distributed system based on the consumed data.

[0118] In the embodiments of this application, there will be a certain time difference between the completion of the migration of existing data (first data) and the start of the synchronization of incremental data (second data). Due to the 24 / 7 nature of the business, there will be a certain amount of data backlog during this time difference. Therefore, a relatively safe synchronization and storage period needs to be set before the first incremental data migration. This period is usually based on the time required for the migration of existing data (e.g., 2 hours). After the subsequent incremental synchronization starts, a message monitoring thread is created synchronously to monitor the processing progress of incremental migration in real time. If there is a message delay, an alarm is issued in time and the user terminal is notified.

[0119] Based on the above database migration method, this application also provides a database migration apparatus. The following will be combined with... Figure 4 The device is described in detail.

[0120] Figure 4 A schematic block diagram of a database migration apparatus according to an embodiment of this application is shown.

[0121] like Figure 4 As shown, the database migration device 400 of this embodiment includes an acquisition module 410, a first partitioning module 420, a second partitioning module 430, and a migration module 440.

[0122] The acquisition module 410 is used to acquire the first data of the database to be migrated. In one embodiment, the acquisition module 410 can be used to perform the operation S210 described above, which will not be repeated here.

[0123] The first partitioning module 420 is used to partition the first data according to a first partitioning rule to obtain multiple first shards, wherein the first partitioning rule matches the architecture of the database to be migrated. In one embodiment, the first partitioning module 420 can be used to perform the operation S220 described above, which will not be repeated here.

[0124] The second partitioning module 430 is used to partition the data included in each first partition according to a second partitioning rule to obtain multiple second partitions, wherein the second partitioning rule matches the architecture of the distributed database. In one embodiment, the second partitioning module 430 can be used to perform the operation S230 described above, which will not be repeated here.

[0125] The migration module 440 is used to migrate the data included in each second shard to the corresponding node of the distributed database. In one embodiment, the migration module 440 can be used to perform the operation S240 described above, which will not be repeated here.

[0126] According to embodiments of this application, any multiple modules among the acquisition module 410, the first partitioning module 420, the second partitioning module 430, and the migration module 440 can be merged into one module, or any one of these modules can be split into multiple modules. Alternatively, at least some of the functions of one or more of these modules can be combined with at least some of the functions of other modules and implemented in one module. According to embodiments of this application, at least one of the acquisition module 410, the first partitioning module 420, the second partitioning module 430, and the migration module 440 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging the circuitry, or implemented in any one of the three implementation methods of software, hardware, and firmware, or in a suitable combination of any of these. Alternatively, at least one of the acquisition module 410, the first partitioning module 420, the second partitioning module 430, and the migration module 440 may be implemented at least partially as a computer program module, which can perform corresponding functions when the computer program module is run.

[0127] In some embodiments, the first partitioning module includes a first partitioning submodule, which is used to perform hash partitioning on the first data according to the key-value information of the first data to obtain multiple first partitions.

[0128] In some embodiments, the second partitioning module includes a second partitioning submodule, which is used to partition the data included in each first partition according to the geographical coordinates of each node of the distributed database and the data characteristics of the first data, to obtain multiple second partitions.

[0129] In some embodiments, the second partitioning submodule includes: a generation unit, configured to generate data identifiers for the data included in each first partition based on the geographical coordinates of each node in the distributed database and the data characteristics of the first data; and a partitioning unit, configured to partition the data included in each first partition based on the data identifiers to obtain multiple second partitions.

[0130] In some embodiments, the apparatus further includes: a first processing module for acquiring second data, wherein the second data is data generated by the database to be migrated during the first data migration process; a second processing module for dividing the second data according to a first partitioning rule to obtain multiple third shards; a third processing module for dividing the data included in each third shard according to the second partitioning rule to obtain multiple fourth shards; and a fourth processing module for migrating the data included in each fourth shard to the corresponding node of the distributed database.

[0131] In some embodiments, the first processing module includes: an acquisition submodule, configured to acquire second data based on relevant table operations of the database to be migrated during the first data migration process.

[0132] In some embodiments, the apparatus further includes a control module, configured to control the migration of second data to the distributed database after waiting for a preset time in response to the migration of first data to the distributed database.

[0133] Figure 5 A block diagram schematically illustrates an electronic device suitable for implementing a database migration method according to an embodiment of this application.

[0134] like Figure 5As shown, an electronic device 500 according to an embodiment of this application includes a processor 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage portion 508 into a random access memory (RAM) 503. The processor 501 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 501 may also include onboard memory for caching purposes. The processor 501 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of this application.

[0135] RAM 503 stores various programs and data required for the operation of electronic device 500. Processor 501, ROM 502, and RAM 503 are interconnected via bus 504. Processor 501 executes various operations of the method flow according to embodiments of this application by executing programs in ROM 502 and / or RAM 503. It should be noted that programs may also be stored in one or more memories other than ROM 502 and RAM 503. Processor 501 may also execute various operations of the method flow according to embodiments of this application by executing programs stored in one or more memories.

[0136] According to embodiments of this application, the electronic device 500 may further include an input / output (I / O) interface 505, which is also connected to a bus 504. The electronic device 500 may also include one or more of the following components connected to the input / output (I / O) interface 505: an input section 506 including a keyboard, mouse, etc.; an output section 507 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the input / output (I / O) interface 505 as needed. A removable medium 511, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 510 as needed so that computer programs read from it can be installed into the storage section 508 as needed.

[0137] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of this application.

[0138] According to embodiments of this application, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this application, the computer-readable storage medium may include ROM 502 and / or RAM 503 and / or one or more memories other than ROM 502 and RAM 503 described above.

[0139] Embodiments of this application also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code is used to enable the computer system to implement the database migration method provided in the embodiments of this application.

[0140] When the computer program is executed by the processor 501, it performs the functions defined in the system / apparatus of this application embodiment. According to the embodiments of this application, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0141] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 509, and / or installed from a removable medium 511. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.

[0142] In such an embodiment, the computer program can be downloaded and installed from a network via communication section 509, and / or installed from removable medium 511. When the computer program is executed by processor 501, it performs the functions defined in the system of this application embodiment. According to embodiments of this application, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0143] According to embodiments of this application, program code for executing the computer programs provided in the embodiments of this application can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C", or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0144] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0145] Those skilled in the art will understand that the features described in the various embodiments of this application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, the features described in the various embodiments of this application can be combined and / or combined in various ways without departing from the spirit and teachings of this application. All such combinations and / or combinations fall within the scope of this application.

Claims

1. A database migration method, characterized in that, The method includes: Retrieve the first data from the database to be migrated; The first data is divided according to the first partitioning rule to obtain multiple first partitions, wherein the first partitioning rule matches the architecture of the database to be migrated; The data included in each first partition is divided according to the second partitioning rule to obtain multiple second partitions, wherein the second partitioning rule matches the architecture of the distributed database. The data included in each second shard is migrated to the corresponding node of the distributed database.

2. The method according to claim 1, characterized in that, The first data is divided according to the first partitioning rule to obtain multiple first fragments, including: The first data is hash-sharded based on the key-value information of the first data to obtain the plurality of first shards.

3. The method according to claim 1, characterized in that, The data in each first partition is divided according to the second partitioning rule to obtain multiple second partitions, including: Based on the geographical coordinates of each node in the distributed database and the data characteristics of the first data, the data included in each first partition is divided to obtain the plurality of second partitions.

4. The method according to claim 3, characterized in that, The process involves dividing the data in each first partition into multiple second partitions based on the geographical coordinates of each node in the distributed database and the data characteristics of the first data, including: Based on the geographical coordinates of each node in the distributed database and the data characteristics of the first data, a data identifier is generated for the data included in each first shard. The data included in each first segment is divided according to the data identifier to obtain the plurality of second segments.

5. The method according to claim 1, characterized in that, The method further includes: Obtain second data, wherein the second data is data generated by the database to be migrated during the migration of the first data; The second data is divided according to the first partitioning rule to obtain multiple third fragments; The data included in each third partition is divided according to the second partitioning rule to obtain multiple fourth partitions; The data included in each fourth shard is migrated to the corresponding node of the distributed database.

6. The method according to claim 5, characterized in that, The acquisition of the second data includes: The second data is obtained based on the relevant table operations of the database to be migrated during the first data migration process.

7. The method according to claim 5, characterized in that, The method further includes: In response to the migration of the first data to the distributed database, after waiting for a preset time, the system controls the migration of the second data to the distributed database.

8. A database migration device, characterized in that, The device includes: The acquisition module is used to acquire the first data from the database to be migrated. The first partitioning module is used to partition the first data according to the first partitioning rule to obtain multiple first partitions, wherein the first partitioning rule matches the architecture of the database to be migrated. The second partitioning module is used to partition the data included in each first partition according to the second partitioning rule to obtain multiple second partitions, wherein the second partitioning rule is matched with the architecture of the distributed database. The migration module is used to migrate the data included in each second shard to the corresponding node of the distributed database.

9. An electronic device, comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 7.

11. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 7.