Online capacity expansion method and cloud database system

By performing full backup and recovery in the cloud database system, combined with the methods of determining the starting site of the binlog file and replaying incremental data, the problem of data migration time in the existing technology is solved, and faster online expansion is achieved.

CN119961350APending Publication Date: 2025-05-09HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410391648.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-08
Filing Date
2024-03-30
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

When the data shard migration is carried out in the prior art, as the amount of data increases, the time required for online expansion is too long, affecting business continuity.

Method used

By performing full backup and recovery in the cloud database system, data shards are migrated, and appropriate starting sites are determined using binlog files, incremental data is extracted and played back to shorten the data migration time.

Benefits of technology

This method significantly shortens the time required for data migration, reduces the time required for online capacity expansion, and avoids the problem of increasing data replay due to XA transaction rollback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961350A_ABST
    Figure CN119961350A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an online capacity expansion method and a cloud database system, and the method comprises the steps: carrying out the full backup of one or more first data fragments of a first database sub-library, and recovering the one or more first data fragments subjected to the full backup into a second database sub-library, the second database sub-library is a new database sub-library for the logic library during capacity expansion. And determining a first start site of a binlog file of the first database sub-library according to the backup time of the one or more first data fragments. And extracting first incremental data of the one or more first data fragments behind the first start site, and replaying the first incremental data to the one or more first data fragments in the second database sub-library.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on November 8, 2023, with application number 202311493265.5 and invention name “A data processing method, device and computing device cluster”, all contents of which are incorporated by reference in this application. Technical Field

[0002] The present application relates to the field of computers, and in particular to an online capacity expansion method and a cloud database system. Background Art

[0003] Cloud database is a stable, reliable and elastically scalable online database service. This database is deployed in a virtual computing environment and managed through a unified management system, which greatly reduces the cost of human operation and maintenance. The sharding middleware is an online database middleware service that focuses on solving the distributed expansion problem of databases, breaking through the capacity and performance bottlenecks of traditional databases, and realizing high concurrent access to massive data. The sharding middleware does not save data itself. All data is stored in multiple data shards of multiple cloud database instances on the backend. In the process of continuous development of customer business, when the performance or storage of the cloud database reaches a bottleneck, some of the data shards need to be migrated online to the newly added cloud database. After migrating the data of the source data shards, the new data generated by the customer in the source data shards needs to be replayed to the new data shards in real time. The sharding middleware ensures that the entire process does not affect the customer's business and does not require the customer to stop, so the above process is also called online capacity expansion.

[0004] During the migration of existing data shards, select SQL is used to extract the full data of the data shards to be migrated on the source side, and then insert SQL is used to write the extracted full data into the newly added cloud database.

[0005] However, when the amount of data to be migrated is too large, the online expansion of the existing technology takes too long. Summary of the invention

[0006] The present application provides an online capacity expansion method and a cloud database system, which are used to shorten the time required for online capacity expansion.

[0007] The first aspect of the present application provides an online capacity expansion method:

[0008] The method is applied to a cloud database system, which includes a cloud database management system and a sub-library and sub-table middleware cluster. The cloud database system provides a logical database to a user, and the sub-library and sub-table middleware cluster is used to manage multiple cloud database instances of the same logical database. Multiple cloud database instances store multiple database sub-libraries of the logical database, each database sub-library includes at least one data shard, and the multiple database sub-libraries include a first database sub-library. The method includes: the cloud database management system performs a full backup of one or more first data shards of the first database sub-library, and the cloud database management system restores the one or more first data shards of the full backup to a second database sub-library, and the second database sub-library is a database sub-library newly added to the logical library during capacity expansion. The sub-library and sub-table middleware cluster determines the first starting point of the binlog file of the first database sub-library according to the backup time of the one or more first data shards, the sub-library and sub-table middleware cluster extracts the first incremental data of the one or more first data shards after the first starting point, and the sub-library and sub-table middleware cluster plays the first incremental data back to the one or more first data shards in the second database sub-library.

[0009] In this application, the essence of full backup of data shards is to upload the data shards to a storage space, and the essence of restoring data shards is to download the data shards from the storage space to the newly added database sub-library. Compared with SQL operations, this method can greatly shorten the time required for data migration, thereby shortening the time required for online expansion.

[0010] In a possible implementation, determining the first starting point of the binlog file of the first database shard according to the backup time of one or more first data shards is specifically:

[0011] Determine a target binlog file in the first database sub-library whose generation time is before the backup time of one or more first data shards, and determine whether there is a target type site in the target binlog file, the target type site is the end point of the XA transaction, or the end point of the data definition language DDL operation. If it exists, determine the first starting site according to the target type site.

[0012] In the present application, since the end point of the DDL operation is not within the XA transaction, the first starting point is the end point of the XA transaction or the end point of the DDL operation, ensuring that the first starting point is not within the XA transaction, thereby avoiding the problem of increased playback data volume due to the rollback of the XA transaction.

[0013] In a possible implementation, determining the first starting site according to the target type site is specifically:

[0014] The last target type position in the target binlog file is determined as the first starting position.

[0015] In the present application, the position of the first starting point is moved back as much as possible, thereby further reducing the amount of data that needs to be played back and shortening the time required for online expansion.

[0016] In a possible implementation, if the target type site does not exist, the starting point of the target binlog file is determined as the first starting site.

[0017] In this application, since the starting point of the target binlog file is generally not located in an XA transaction, the problem of an increase in the amount of playback data due to the rollback result of the XA transaction is avoided.

[0018] In a possible implementation, the multiple database sub-libraries further include a third database sub-library, and the method further includes:

[0019] The cloud database management system performs a full backup of one or more second data shards of the third database sub-library, wherein the backup time of the one or more second data shards is different from the backup time of the one or more first data shards. The cloud database management system restores the one or more second data shards of the full backup to the second database sub-library, and the second database sub-library is a database sub-library newly added to the logical library during capacity expansion. The sub-library and table middleware cluster determines the second starting point of the binlog file of the third database sub-library based on the backup time of the one or more second data shards. The sub-library and table middleware cluster extracts the second incremental data of the one or more second data shards after the second starting point, and the sub-library and table middleware cluster replays the second incremental data to the one or more second data shards in the second database sub-library.

[0020] The second aspect of the present application provides a cloud database system:

[0021] The cloud database system includes a cloud database management system and a sub-library and sub-table middleware cluster. The cloud database system provides a logical database to users. The sub-library and sub-table middleware cluster is used to manage multiple cloud database instances of the same logical database. Multiple cloud database instances store multiple database sub-libraries of the logical database, each database sub-library includes at least one data shard, and the multiple database sub-libraries include a first database sub-library.

[0022] The cloud database management system is used to perform a full backup of one or more first data shards of a first database sub-library.

[0023] The cloud database management system is also used to restore one or more first data shards of the full backup to the second database shard, and the second database shard is a database shard newly added to the logical library during capacity expansion.

[0024] The sharding middleware cluster is used to determine the first starting point of the binlog file of the first database shard according to the backup time of one or more first data shards.

[0025] The database and table sharding middleware cluster is also used to extract the first incremental data of one or more first data shards after the first starting point.

[0026] The database and table sharding middleware cluster is also used to replay the first incremental data to one or more first data shards in the second database shard.

[0027] In this application, data shards are migrated by means of backup and recovery, which greatly shortens the time required for data migration and thus shortens the time required for online capacity expansion.

[0028] In one possible implementation,

[0029] The sharding middleware cluster is specifically used to determine a target binlog file in the first database shard whose generation time is before the backup time of one or more first data shards.

[0030] The sharding middleware cluster is specifically used to determine whether there is a target type site in the target binlog file. The target type site is the end point of the XA transaction or the end point of the data definition language DDL operation.

[0031] The sub-library and sub-table middleware cluster is specifically used to determine the first starting point according to the target type site if there is a target type site.

[0032] In the present application, since the end point of the DDL operation is not within the XA transaction, the first starting point is the end point of the XA transaction or the end point of the DDL operation, ensuring that the first starting point is not within the XA transaction, thereby avoiding the problem of increased playback data volume due to the rollback of the XA transaction.

[0033] In one possible implementation,

[0034] The database and table sharding middleware cluster is specifically used to determine the last target type position in the target binlog file as the first starting position.

[0035] In the present application, the position of the first starting point is moved back as much as possible, thereby further reducing the amount of data that needs to be played back and shortening the time required for online expansion.

[0036] In one possible implementation,

[0037] The database and table sharding middleware cluster is specifically used to determine the starting point of the target binlog file as the first starting point if the target type site does not exist.

[0038] In this application, since the starting point of the target binlog file is generally not located in an XA transaction, the problem of an increase in the amount of playback data due to the rollback result of the XA transaction is avoided.

[0039] In a possible implementation, the multiple database sub-libraries further include a third database sub-library;

[0040] The cloud database management system is also used to fully back up one or more second data shards of the third database sub-library, wherein the backup time of the one or more second data shards is different from the backup time of the one or more first data shards.

[0041] The cloud database management system is also used to restore one or more second data shards of the full backup to the second database sub-library, and the second database sub-library is a database sub-library newly added to the logical library during capacity expansion.

[0042] The database and table sharding middleware cluster is also used to determine the second starting point of the binlog file of the third database shard according to the backup time of one or more second data shards.

[0043] The database and table sharding middleware cluster is also used to extract second incremental data of one or more second data shards after the second starting point.

[0044] The sharded database and table middleware cluster is also used to replay the second incremental data to one or more second data shards in the second database shard.

[0045] In a third aspect, the present application provides a computing device cluster, including at least one computing device, each computing device including a processor and a memory. The processor of at least one computing device is used to execute instructions stored in the memory of at least one computing device, so that the computing device cluster executes the method in the first aspect as described above.

[0046] A fourth aspect of the present application provides a computer program product comprising instructions, which, when executed by a computing device cluster, causes the computing device cluster to execute the method in the aforementioned first aspect.

[0047] A fifth aspect of the present application provides a computer-readable storage medium, comprising computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster executes the method in the aforementioned first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1a This is a flowchart of the existing online capacity expansion;

[0049] Figure 1b A schematic diagram of the system architecture used in this application;

[0050] Figure 2 A schematic diagram of a flow chart of the online capacity expansion method in this application;

[0051] Figure 3 A schematic diagram of full backup of data shards in this application;

[0052] Figure 4 A schematic diagram of recovering data shards in this application;

[0053] Figure 5 This is a schematic diagram for determining the target binlog file in this application;

[0054] Figure 6 A schematic diagram for determining the starting point in this application;

[0055] Figure 7 Another schematic diagram of full backup of data shards in this application;

[0056] Figure 8 Another schematic diagram of restoring data shards in this application;

[0057] Fig. 9 A schematic diagram of a cloud database system in this application;

[0058] Fig.10 A schematic diagram of a computing device in the present application;

[0059] Fig.11 A schematic diagram of a computing device cluster in this application;

[0060] Fig.12 Another schematic diagram of a computing device cluster in the present application. DETAILED DESCRIPTION

[0061] The following describes the embodiments of the present application in conjunction with the accompanying drawings. Obviously, the described embodiments are only embodiments of a part of the present application, rather than all embodiments. It is known to those skilled in the art that with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0062] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0063] In order to facilitate the understanding of this application, the following first introduces the relevant concepts involved in this application:

[0064] XA transaction: In order to achieve a distributed transaction protocol that implements distributed consistency, the start and end of an XA transaction will record corresponding events in the binlog file. XA transactions are implemented using a two-phase commit (2PC). The first phase is Prepare. The transaction manager (TM) sends prepare instructions to all resource managers (RMs). RMs perform data modifications and log records, and return messages to TM that can be submitted or not. When TM receives prepared messages from all RMs, it notifies all transactions to commit and enters the second phase. The second phase is Commit. TM receives the prepare results from all RMs. If any RM returns an uncommittable or timed-out state, it sends a Rollback command to all RMs. If all RMs return a committable state, it sends a Commit command to all RMs to complete the transaction operation.

[0065] Data definition language (DDL) operation: DDL is a part of SQL language, which is used to define and modify the structure of the database, including creating tables, deleting tables, modifying table structures, etc. DDL operations are related to the structure of the database and involve data storage and access efficiency.

[0066] Binlog files: The cloud database generates incremental data every day. If the full data of the cloud database is backed up once a day, it will greatly increase the unnecessary burden. Therefore, the cloud database will write the incremental data into a series of binlog files, such as mysql-bin.000001, mysql-bin.000002, mysql-bin.000003, etc. The location of the incremental data can be tracked by the location (file name + file location).

[0067] Sub-library and table-splitting middleware example: Sub-library and table-splitting middleware provides users with a logical database whose data is stored on multiple physical storage instances. The physical library on each storage instance is called a sub-library.

[0068] See also Figure 1a Taking the migration of data shard 1 of cloud database 1 to cloud database 2 as an example, the existing online capacity expansion method is mainly divided into the following steps:

[0069] 1. First, the database and table sharding middleware cluster determines the appropriate starting point in the binlog file corresponding to cloud database 1.

[0070] 2. The sharded database and table middleware cluster extracts the full data of data shard 1 that needs to be migrated on the source side through select SQL.

[0071] 3. The sharded database and table middleware cluster writes the full amount of data of the data shard 1 to be migrated into the newly added cloud database 2 through insert SQL.

[0072] 4. The database and table sharding middleware cluster extracts the incremental data after the starting point of data shard 1.

[0073] 5. The sharded database and table middleware cluster plays the incremental data back to data shard 1 in cloud database 2.

[0074] 6. After the data is equalized, suspend customer services and complete the route switching.

[0075] In the above technical solution, the migration time increases with the increase of data volume. In the case of massive data, the full data migration time may take several days. In addition, the incremental data generated during the longer migration time also increases accordingly, and more time is required to replay the incremental data, which further prolongs the overall change time.

[0076] This application can be applied to Figure 1b In the cloud database system shown, the cloud database system includes a cloud database management system, a sub-library and sub-table middleware cluster, and a sub-library and sub-table middleware management system, wherein the sub-library and sub-table middleware cluster is used to manage multiple cloud database instances of the same logical database, and multiple cloud database instances store multiple database sub-libraries of the logical database (cloud database 1, cloud database 2, etc. in the figure), and each database sub-library includes at least one data shard. The sub-library and sub-table middleware management system can manage the sub-library and sub-table middleware cluster.

[0077] See also Figure 2 , the following is an introduction to a process of the online capacity expansion method of this application:

[0078] 201. The cloud database management system performs a full backup of one or more first data shards of a first database sub-library;

[0079] See also Figure 3 , assuming that the logical database includes three database sub-databases, namely, cloud database 1, cloud database 2, and cloud database 3 newly added for online expansion, wherein cloud database 1 includes data shard 1, data shard 2, and data shard 3, and cloud database 2 includes data shard 4, data shard 5, and data shard 6. If data shard 2 and data shard 3 need to be migrated to cloud database 3, the sub-database and sub-table middleware management system will send backup instruction information to the cloud database management system, wherein the backup instruction information is used to instruct the cloud database management system to perform a full backup of data shard 2 and data shard 3 in cloud database 1, and the sub-database and sub-table middleware management system will record the time when the backup instruction information is sent, which is recorded as backup time 1. After receiving the above backup instruction information, the cloud database management system will perform a full backup of data shard 2 and data shard 3.

[0080] 202. The cloud database management system restores the one or more first data shards of the full backup to the second database shard, where the second database shard is a database shard added to the logic library during capacity expansion.

[0081] See also Figure 4 After the full backup of data shard 2 and data shard 3 is completed, the cloud database management system sends a backup completion response to the sub-library and sub-table middleware management system. After receiving the response, the sub-library and sub-table middleware management system sends a recovery instruction message to the cloud database management system. The recovery instruction message is used to instruct the cloud database management system to restore the fully backed-up data shard 2 and data shard 3 to cloud database 3. After receiving the recovery instruction message, the cloud database management system performs a recovery operation to restore the fully backed-up data shard 2 and data shard 3 to cloud database 3.

[0082] 203. The sub-library and sub-table middleware cluster determines the first starting point of the binlog file of the first database sub-library according to the backup time of one or more first data shards;

[0083] See also Figure 5 After restoring the fully backed-up data shard 2 and data shard 3 to the cloud database 3, the cloud database management system sends a recovery completion response to the sub-library and sub-table middleware management system. After receiving the response, the sub-library and sub-table middleware management system sends an extraction indication message to the sub-library and sub-table middleware cluster, which carries the above-mentioned backup time 1.

[0084] The sharding middleware cluster determines a binlog file whose generation time is before the backup time 1 in the multiple binlog files corresponding to the cloud database 1. The sharding middleware cluster determines whether there is a target site in the binlog file. The target site can be the end point of the XA transaction or the end point of the DDL operation. If there is a target site, the sharding middleware cluster determines the target site as the starting site. Since the end point of the DDL operation is generally not located in the XA transaction, the starting site determined by the above method generally does not fall within the XA transaction, thereby avoiding the problem of increased playback data when the result of the XA transaction is a rollback. Optionally, if there are multiple target sites in the above binlog file, the sharding middleware cluster determines the last target site as the starting site, thereby moving the starting site back as much as possible to reduce the amount of data that needs to be replayed. If the above target site does not exist in the binlog file, the starting point of the binlog file is directly determined as the starting site.

[0085] 204. The database and table sharding middleware cluster extracts the first incremental data of one or more first data shards after the first starting point;

[0086] See also Figure 6 After determining the starting point, the database and table sharding middleware cluster obtains the data of multiple binlog files of cloud database 1 after the starting point, and filters the data to retain only the incremental data of data shard 2 and data shard 3.

[0087] 205. The sub-library and table middleware cluster plays back the first incremental data to one or more first data shards in the second database sub-library.

[0088] The sharding middleware cluster plays back the incremental data of data shard 2 and data shard 3 to data shard 2 and data shard 3 of cloud database 3.

[0089] In addition to migrating data shards from cloud database 1 to cloud database 3, the cloud database system can also migrate data shards from cloud database 2 to cloud database 3. The following is a detailed description:

[0090] The sub-library and sub-table middleware management system sends backup instruction information to the cloud database management system, wherein the backup instruction information is used to instruct the cloud database management system to perform a full backup of data shard 4 and data shard 5 in cloud database 2. The sub-library and sub-table middleware management system records the time when the backup instruction information is sent, and the time is recorded as backup time 2. After receiving the above backup instruction information, the cloud database management system performs a full backup of data shard 2 and data shard 3. It should be noted that the above backup time 2 is not the same as the backup time 1.

[0091] After the cloud database management system completes the full backup of data shard 4 and data shard 5, it sends a backup completion response to the sub-library and sub-table middleware management system. After receiving the response, the sub-library and sub-table middleware management system sends a recovery instruction message to the cloud database management system, and the recovery instruction message is used to instruct the cloud database management system to restore the full-backed data shard 4 and data shard 5 to the cloud database 3. After receiving the recovery instruction message, the cloud database management system performs a recovery operation and restores the full-backed data shard 4 and data shard 5 to the cloud database 3.

[0092] After restoring the fully backed-up data shards 4 and 5 to the cloud database 3, the cloud database management system sends a recovery completion response to the sub-library and table middleware management system. After receiving the response, the sub-library and table middleware management system sends extraction indication information to the sub-library and table middleware cluster, which carries the above-mentioned backup time 2.

[0093] The sub-library and sub-table middleware cluster determines a binlog file whose generation time is before the backup time 2 from the multiple binlog files corresponding to the cloud database 2. The sub-library and sub-table middleware cluster determines whether there is a target site in the binlog file. If there is a target site, the sub-library and sub-table middleware cluster determines the target site as the starting site. If there are multiple target sites in the above binlog file, the sub-library and sub-table middleware cluster determines the last target site as the starting site. If the above target site does not exist in the binlog file, the starting point of the binlog file is directly determined as the starting site. After determining the starting site, the sub-library and sub-table middleware cluster obtains the data of the multiple binlog files of the cloud database 2 after the starting site, and filters these data, retaining only the incremental data of data shard 4 and data shard 5. Finally, the sub-library and sub-table middleware cluster replays the incremental data of data shard 4 and data shard 5 to data shard 4 and data shard 5 of the cloud database 3.

[0094] After completing the above operations, the database and table sharding middleware node suspends the business, performs route switching and resumes the business.

[0095] In this application, the data shards are migrated by means of backup and recovery, which greatly shortens the time required for data migration, thereby shortening the time required for online expansion. Based on the appropriate starting point, the amount of data required for playback can be reduced, further shortening the time required for online expansion.

[0096] Another process of the online capacity expansion method of this application is introduced below:

[0097] A01. The cloud database management system performs a full backup of one or more first data shards of the first database sub-library;

[0098] See also Figure 7 , assuming that the logical database includes three database sub-databases, namely, cloud database 1, cloud database 2, and cloud database 3 newly added for online expansion, wherein cloud database 1 includes data shard 1, data shard 2, and data shard 3, and cloud database 2 includes data shard 4, data shard 5, and data shard 6. If it is necessary to migrate data shard 2 and data shard 3 to cloud database 3, the sub-database and sub-table middleware management system will send backup instruction information to the cloud database management system, wherein the backup instruction information is used to instruct the cloud database management system to fully back up data shard 2 and data shard 3 in cloud database 1, and the sub-database and sub-table middleware management system will record the time when the backup instruction information is sent, and the time is recorded as backup time 1. Different from the above-mentioned embodiment, after receiving the above-mentioned backup instruction information, the cloud database management system will fully back up all data shards in cloud database 1.

[0099] A02. The cloud database management system restores one or more first data shards of the full backup to the second database sub-library, where the second database sub-library is a database sub-library newly added to the logical library during capacity expansion;

[0100] See also Figure 8 After the cloud database management system performs a full backup of all data shards of cloud database 1, it sends a backup completion response to the sub-library and sub-table middleware management system. After receiving the response, the sub-library and sub-table middleware management system sends a recovery instruction message to the cloud database management system, and the recovery instruction message is used to instruct the cloud database management system to restore all data shards of the full backup to cloud database 3. After receiving the recovery instruction message, the cloud database management system performs a recovery operation to restore all data shards of the full backup to cloud database 3, and then deletes the other data shards in cloud database 3 except data shard 2 and data shard 3.

[0101] A03. The database and table sharding middleware cluster determines the first starting point of the binlog file of the first database shard according to the backup time of one or more first data shards;

[0102] Similarly, after completing the above operations, the cloud database management system sends a recovery completion response to the sub-library and table middleware management system. After receiving the response, the sub-library and table middleware management system sends extraction indication information to the sub-library and table middleware cluster, which carries the above-mentioned backup time 1.

[0103] The sharding middleware cluster determines a binlog file whose generation time is before the backup time 1 in the multiple binlog files corresponding to the cloud database 1. The sharding middleware cluster determines whether there is a target site in the binlog file. The target site can be the end point of the XA transaction or the end point of the DDL operation. If there is a target site, the sharding middleware cluster determines the target site as the starting site. Since the end point of the DDL operation is generally not located in the XA transaction, the starting site determined by the above method generally does not fall within the XA transaction, thereby avoiding the problem of increased playback data when the result of the XA transaction is a rollback. Optionally, if there are multiple target sites in the above binlog file, the sharding middleware cluster determines the last target site as the starting site, thereby moving the starting site back as much as possible to reduce the amount of data that needs to be replayed. If the above target site does not exist in the binlog file, the starting point of the binlog file is directly determined as the starting site.

[0104] A04. The database and table sharding middleware cluster extracts the first incremental data of one or more first data shards after the first starting point;

[0105] After determining the starting point, the database and table sharding middleware cluster obtains the data of multiple binlog files of cloud database 1 after the starting point, and filters the data to retain only the incremental data of data shard 2 and data shard 3.

[0106] A05. The database and table sharding middleware cluster replays the first incremental data to one or more first data shards in the second database shard.

[0107] The sharding middleware cluster plays back the incremental data of data shard 2 and data shard 3 to data shard 2 and data shard 3 of cloud database 3.

[0108] Another process of the online capacity expansion method of this application is introduced below:

[0109] B01. The cloud database management system performs a full backup of one or more first data shards of the first database sub-library;

[0110] Similarly, assume that the logical database includes three database sub-databases, namely, cloud database 1, cloud database 2, and cloud database 3 newly added for online expansion, wherein cloud database 1 includes data shard 1, data shard 2, and data shard 3, and cloud database 2 includes data shard 4, data shard 5, and data shard 6. If it is necessary to migrate data shard 2 and data shard 3 to cloud database 3, the sub-database and sub-table middleware management system will send backup instruction information to the cloud database management system, wherein the backup instruction information is used to instruct the cloud database management system to perform a full backup of data shard 2 and data shard 3 in cloud database 1, and the sub-database and sub-table middleware management system will record the time when the backup instruction information is sent, and the time is recorded as backup time 1. Different from the above-mentioned embodiment, after receiving the above-mentioned backup instruction information, the cloud database management system will perform a full backup of all data shards in cloud database 1.

[0111] B02. The cloud database management system restores one or more first data shards of the full backup to the second database sub-library, where the second database sub-library is a database sub-library newly added to the logical library during capacity expansion;

[0112] After performing a full backup of all data shards of cloud database 1, the cloud database management system sends a backup completion response to the sub-library and sub-table middleware management system. After receiving the response, the sub-library and sub-table middleware management system sends a recovery instruction message to the cloud database management system, and the recovery instruction message is used to instruct the cloud database management system to restore the full backup data shard 2 and data shard 3 to cloud database 3. After receiving the recovery instruction message, the cloud database management system performs a recovery operation and restores the full backup data shard 2 and data shard 3 to cloud database 3.

[0113] B03. The sub-library and sub-table middleware cluster determines the first starting point of the binlog file of the first database sub-library according to the backup time of one or more first data shards;

[0114] Similarly, after completing the above operations, the cloud database management system sends a recovery completion response to the sub-library and table middleware management system. After receiving the response, the sub-library and table middleware management system sends extraction indication information to the sub-library and table middleware cluster, which carries the above-mentioned backup time 1.

[0115] The sharding middleware cluster determines a binlog file whose generation time is before the backup time 1 in the multiple binlog files corresponding to the cloud database 1. The sharding middleware cluster determines whether there is a target site in the binlog file. The target site can be the end point of the XA transaction or the end point of the DDL operation. If there is a target site, the sharding middleware cluster determines the target site as the starting site. Since the end point of the DDL operation is generally not located in the XA transaction, the starting site determined by the above method generally does not fall within the XA transaction, thereby avoiding the problem of increased playback data when the result of the XA transaction is a rollback. Optionally, if there are multiple target sites in the above binlog file, the sharding middleware cluster determines the last target site as the starting site, thereby moving the starting site back as much as possible to reduce the amount of data that needs to be replayed. If the above target site does not exist in the binlog file, the starting point of the binlog file is directly determined as the starting site.

[0116] B04. The database and table sharding middleware cluster extracts the first incremental data after the first starting point of one or more first data shards;

[0117] After determining the starting point, the database and table sharding middleware cluster obtains the data of multiple binlog files of cloud database 1 after the starting point, and filters the data to retain only the incremental data of data shard 2 and data shard 3.

[0118] B05. The sub-library and table middleware cluster plays back the first incremental data to one or more first data shards in the second database sub-library.

[0119] The sharding middleware cluster plays back the incremental data of data shard 2 and data shard 3 to data shard 2 and data shard 3 of cloud database 3.

[0120] The method in this application is introduced above. The cloud database system in this application is introduced below:

[0121] The present application also provides a cloud database system, such as Fig. 9 As shown, it includes a cloud database management system and a database and table sharding middleware cluster. The database and table sharding middleware cluster is used to manage multiple cloud database instances of the same logical database. Multiple database shards of the logical database are stored on the multiple cloud database instances. Each database shard includes at least one data shard. The multiple database shards include a first database shard.

[0122] The cloud database management system is used to perform a full backup of one or more first data shards of a first database sub-library.

[0123] The cloud database management system is also used to restore one or more first data shards of the full backup to the second database shard, and the second database shard is a database shard newly added to the logical library during capacity expansion.

[0124] The sharding middleware cluster is used to determine the first starting point of the binlog file of the first database shard according to the backup time of one or more first data shards.

[0125] The database and table sharding middleware cluster is also used to extract the first incremental data of one or more first data shards after the first starting point.

[0126] The database and table sharding middleware cluster is also used to replay the first incremental data to one or more first data shards in the second database shard.

[0127] The cloud database management system and the sub-library and sub-table middleware cluster can be implemented by software or hardware. As an example, the implementation of the cloud database management system is introduced below. Similarly, the implementation of the sub-library and sub-table middleware cluster can refer to the implementation of the cloud database management system.

[0128] As an example of a software functional unit, a cloud database management system may include code running on a computing instance. The computing instance may be at least one of a physical host (computing device), a virtual machine, a container, and other computing devices. Furthermore, the computing device may be one or more. For example, a cloud database management system may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the application may be distributed in the same region or in different regions. The multiple hosts / virtual machines / containers used to run the code may be distributed in the same AZ or in different AZs, each AZ including one data center or multiple data centers with similar geographical locations. Typically, a region may include multiple AZs.

[0129] Similarly, multiple hosts / virtual machines / containers used to run the code can be distributed in the same VPC or in multiple VPCs. Usually, a VPC is set up in a region. For cross-region communication between two VPCs in the same region and between VPCs in different regions, a communication gateway must be set up in each VPC to achieve interconnection between VPCs through the communication gateway.

[0130] As an example of a hardware functional unit, the cloud database management system may include at least one computing device, such as a server, etc. Alternatively, the cloud database management system may also be a device implemented using ASIC or PLD, etc. The PLD may be implemented using CPLD, FPGA, GAL or any combination thereof.

[0131] The multiple computing devices included in the cloud database management system can be distributed in the same region or in different regions. The multiple computing devices included in the cloud database management system can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the cloud database management system can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0132] The present application also provides a computing device 100. Fig.10 As shown, the computing device 100 includes: a bus 102, a processor 104, a memory 106, and a communication interface 108. The processor 104, the memory 106, and the communication interface 108 communicate with each other through the bus 102. The computing device 100 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 100.

[0133] The bus 102 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig.10 The bus 104 is represented by only one line, but does not mean that there is only one bus or one type of bus. The bus 104 may include a path for transmitting information between various components of the computing device 100 (eg, the memory 106, the processor 104, and the communication interface 108).

[0134] The processor 104 may include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0135] The memory 106 may include a volatile memory, such as a random access memory (RAM). The processor 104 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0136] The memory 106 stores executable codes, and the processor 104 executes the executable codes to respectively implement the functions of the aforementioned cloud database management system and the sub-library and sub-table middleware cluster, thereby implementing the online capacity expansion method of the present application. That is, the memory 106 stores instructions for executing the online capacity expansion method of the present application.

[0137] The embodiment of the present application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.

[0138] like Fig.11 As shown, the computing device cluster includes at least one computing device 100. The memory 106 in one or more computing devices 100 in the computing device cluster may store the same instructions for executing the online capacity expansion method of the present application.

[0139] In some possible implementations, the memory 106 of one or more computing devices 100 in the computing device cluster may also store some instructions for executing the online capacity expansion method of the present application. In other words, the combination of one or more computing devices 100 can jointly execute the instructions for executing the online capacity expansion method of the present application.

[0140] It should be noted that the memory 106 in different computing devices 100 in the computing device cluster can store different instructions, which are respectively used to execute part of the functions of the cloud database system. That is, the instructions stored in the memory 106 in different computing devices 100 can implement the functions of one or more devices in the cloud database management system and the sub-library and sub-table middleware cluster.

[0141] In some possible implementations, one or more computing devices in the computing device cluster may be connected via a network, which may be a wide area network or a local area network. Fig.12 A possible implementation is shown. Fig.12 As shown, two computing devices 100A and 100B are connected via a network. Specifically, the network is connected via a communication interface in each computing device. In this type of possible implementation, the memory 106 in the computing device 100A stores instructions for executing the functions of the cloud database management system. At the same time, the memory 106 in the computing device 100B stores instructions for executing the functions of the sub-library and sub-table middleware cluster.

[0142] It should be understood that Fig.12 The functions of the computing device 100A shown in FIG. 1 may also be completed by multiple computing devices 100. Similarly, the functions of the computing device 100B may also be completed by multiple computing devices 100.

[0143] The present application embodiment also provides another computing device cluster. The connection relationship between the computing devices in the computing device cluster can be similar to that of Figure 4 and Figure 5 The connection mode of the computing device cluster is different in that the memory 106 in one or more computing devices 100 in the computing device cluster may store the same instructions for executing the online capacity expansion method.

[0144] In some possible implementations, the memory 106 of one or more computing devices 100 in the computing device cluster may also store some instructions for executing the online capacity expansion method. In other words, the combination of one or more computing devices 100 can jointly execute the instructions for executing the online capacity expansion method.

[0145] The embodiment of the present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the online capacity expansion method.

[0146] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk). The computer-readable storage medium includes instructions that instruct the computing device to execute the online expansion method.

[0147] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.

Claims

1. An online capacity expansion method, characterized in that: The method is applied to a cloud database system, the cloud database system includes a cloud database management system and a sub-library and sub-table middleware cluster, the sub-library and sub-table middleware cluster is used to manage multiple cloud database instances of the same logical database, the multiple cloud database instances store multiple database sub-libraries of the logical database, each database sub-library includes at least one data shard, the multiple database sub-libraries include a first database sub-library, the method includes: The cloud database management system performs a full backup of one or more first data shards of the first database sub-library; The cloud database management system restores the one or more first data shards of the full backup to a second database sub-library, where the second database sub-library is a database sub-library newly added to the logic library during capacity expansion; The sub-library and sub-table middleware cluster determines the first starting point of the binlog file of the first database sub-library according to the backup time of the one or more first data shards; The sub-library and sub-table middleware cluster extracts the first incremental data of the one or more first data shards after the first starting point; The sharded database and table middleware cluster plays the first incremental data back to the one or more first data shards in the second database shard.

2. The method according to claim 1, characterized in that Determining the first starting point of the binlog file of the first database partition according to the backup time of the one or more first data partitions includes: Determine a target binlog file in the first database shard whose generation time is before the backup time of the one or more first data shards; Determine whether there is a target type site in the target binlog file, where the target type site is the end point of the XA transaction or the end point of the data definition language DDL operation; If so, the first starting site is determined according to the target type site.

3. The method according to claim 2, characterized in that Determining the first starting site according to the target type site includes: The last target type position in the target binlog file is determined as the first starting position.

4. The method according to claim 2 or 3, characterized in that: If the target type site does not exist, the starting point of the target binlog file is determined as the first starting site.

5. The method according to any one of claims 1 to 4, characterized in that The plurality of database sub-libraries further include a third database sub-library, and the method further includes: The cloud database management system performs a full backup of the one or more second data shards of the third database sub-library, wherein the backup time of the one or more second data shards is different from the backup time of the one or more first data shards; The cloud database management system restores the one or more second data shards of the full backup to the second database sub-library, where the second database sub-library is a database sub-library newly added to the logic library during capacity expansion; The sub-library and sub-table middleware cluster determines the second starting point of the binlog file of the third database sub-library according to the backup time of the one or more second data shards; The sub-library and sub-table middleware cluster extracts second incremental data of the one or more second data shards after the second starting point; The sub-library and table middleware cluster plays back the second incremental data to the one or more second data shards in the second database sub-library.

6. A cloud database system, characterized in that: It includes a cloud database management system and a sub-library and sub-table middleware cluster, wherein the sub-library and sub-table middleware cluster is used to manage multiple cloud database instances of the same logical database, wherein the multiple cloud database instances store multiple database sub-libraries of the logical database, each database sub-library includes at least one data shard, and the multiple database sub-libraries include a first database sub-library; The cloud database management system is used to perform a full backup of one or more first data shards of the first database sub-library; The cloud database management system is further used to restore the one or more first data shards of the full backup to a second database sub-library, where the second database sub-library is a database sub-library newly added to the logic library during capacity expansion; The sub-library and sub-table middleware cluster is used to determine the first starting point of the binlog file of the first database sub-library according to the backup time of the one or more first data shards; The sub-library and sub-table middleware cluster is further used to extract first incremental data of the one or more first data shards after the first starting point; The sub-library and table middleware cluster is also used to replay the first incremental data to the one or more first data shards in the second database sub-library.

7. The cloud database system according to claim 6, characterized in that: The sub-library and sub-table middleware cluster is specifically used to determine a target binlog file in the first database sub-library whose generation time is before the backup time of the one or more first data shards; The sub-library and sub-table middleware cluster is specifically used to determine whether there is a target type site in the target binlog file, and the target type site is the end point of the XA transaction or the end point of the data definition language DDL operation; The sub-library and sub-table middleware cluster is specifically used to determine the first starting point according to the target type site if the target type site exists.

8. The cloud database system according to claim 7, characterized in that: The sub-library and sub-table middleware cluster is specifically used to determine the last target type site in the target binlog file as the first starting site.

9. The cloud database system according to claim 7 or 8, characterized in that: The sub-library and sub-table middleware cluster is specifically used to determine the starting point of the target binlog file as the first starting point if the target type site does not exist.

10. The cloud database system according to any one of claims 6 to 9, characterized in that: The multiple database sub-libraries also include a third database sub-library; The cloud database management system is further used to perform a full backup of one or more second data shards of the third database sub-library, wherein the backup time of the one or more second data shards is different from the backup time of the one or more first data shards; The cloud database management system is further used to restore the one or more second data shards of the full backup to the second database sub-library, where the second database sub-library is a database sub-library newly added to the logical library during capacity expansion; The sub-library and sub-table middleware cluster is further used to determine the second starting position of the binlog file of the third database sub-library according to the backup time of the one or more second data shards; The sub-library and sub-table middleware cluster is further used to extract second incremental data of the one or more second data shards after the second starting point; The sub-library and table middleware cluster is also used to replay the second incremental data to the one or more second data shards in the second database sub-library.

11. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 5.

12. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device cluster, the computing device cluster is caused to perform the method according to any one of claims 1 to 5.

13. A computer-readable storage medium, characterized in that: The method comprises computer program instructions, and when the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method as claimed in any one of claims 1 to 5.