Data hot backup method and related device

By establishing a data pipeline between the computing device and the storage device in the cloud database and using heartbeat data blocks to maintain the data pipeline connection, the problem of large overhead of hot backup resources in the prior art is solved, and efficient and reliable hot backup is achieved.

CN120029819APending Publication Date: 2025-05-23HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311573339.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-22
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The hot backup method in the existing cloud database needs to occupy resources for backup cloud hard disks, resulting in a large resource overhead.

Method used

By establishing a data pipeline between the computing device and the storage device, and generating a heartbeat data block, sending a heartbeat data block to the storage device according to the heartbeat cycle, reading the backup data from the cloud hard disk, dividing it into multiple data segments, and sending the backup data segment to the storage device through the data pipeline.

Benefits of technology

No need to use backup cloud hard disk to read and write data, reducing the resource overhead of hot backups and improving the reliability of backups.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029819A_ABST
    Figure CN120029819A_ABST
Patent Text Reader

Abstract

The invention provides a data hot backup method, which is applied to the field of cloud databases, and does not use a backup cloud hard disk to store backup data in a hot backup process, so that the resource overhead can be saved. The method comprises the following steps: after receiving a backup instruction input by a user, a computing device establishes a data pipeline between the computing device and a storage device and generates a heartbeat data block in response to the backup instruction, and sends the heartbeat data block to the storage device through the data pipeline according to a heartbeat cycle; after the backup data is read from the cloud hard disk, the backup data is divided into a plurality of data segments, the data segments are sent to the storage device through the data pipeline, and the data pipeline between the computing device and the storage device can be kept unbroken through the heartbeat data block, so that the reliability of hot backup can be improved under the condition that the backup cloud hard disk is not used.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of cloud databases, and in particular to a data hot backup method, a computing device, a storage device, a cloud database system, a computing device cluster, a computer-readable storage medium, and a computer program product. Background Art

[0002] A cloud database is a database deployed in a virtual computing environment, with advantages such as pay-as-you-go, on-demand expansion, high availability, and storage consolidation.

[0003] The current method of hot backup in cloud database is as follows: when the cloud server runs the database instance, when the cloud server has idle resources, read the backup data from cloud hard disk 1, divide the backup data into data slices of fixed size, write the data slices into cloud hard disk 2 (i.e., backup cloud hard disk) in the form of data stream, read the complete data slices from cloud hard disk 2, and send the data slices to the storage device, such as Figure 1 shown.

[0004] The above hot backup method needs to read and write data in the cloud hard disk 2, so it will occupy the resources of the cloud hard disk 2. Summary of the invention

[0005] The present application provides a data hot backup method, which does not require the use of a backup cloud hard disk to read and write data, thereby reducing the resource overhead of hot backup. The present application also provides related devices for implementing the method, such as a computing device, a storage device, a cloud database system, a computing device cluster, a computer-readable storage medium, and a computer program product.

[0006] The first aspect of the present application provides a data hot backup method, which is applied to a computing device of a cloud database system, the computing device is in a state of running a database instance, after the computing device receives a backup instruction input by a user, a data pipeline is established between the computing device and a storage device in response to the backup instruction and a heartbeat data block is generated, and the heartbeat data block is sent to the storage device through the data pipeline according to the heartbeat cycle; after the computing device reads the backup data from the cloud hard disk, the backup data is divided into multiple data segments, and the data segments of the backup data are sent to the storage device through the data pipeline. The backup data includes at least one of a redo log, a system tablespace file, a clustered index data table, or a non-clustered index data table.

[0007] A heartbeat data block refers to a data block sent according to a heartbeat cycle, and the heartbeat cycle is less than the connection duration threshold of the data pipeline. In this implementation, sending heartbeat data blocks according to the heartbeat cycle can keep the data pipeline between the computing device and the storage device from being interrupted, and can improve the reliability of hot backup without using a backup cloud hard disk (i.e., the backup data is not stored on the disk), thereby reducing the resource overhead of hot backup.

[0008] In a possible implementation, the data hot backup method of the present application also includes: the computing device generates a backup tool log including the backup status and the data volume of the backup data; when it is detected that the backup status is the backup completion status and the data volume of the backup data is consistent with the preset backup data volume, the data pipeline is closed; when it is detected that the backup status is the backup completion status and the data volume of the backup data is inconsistent with the preset backup data volume, the computing device re-reads the backup data from the cloud hard disk; divides the backup data into multiple data segments; and sends the data segments of the backup data to the storage device through the data pipeline. In this way, the backup progress can be obtained in time, and the data pipeline can be quickly closed after the backup is completed to save network bandwidth overhead. When it is detected that the backup status is the backup completion status and the data volume of the backup data is inconsistent with the preset backup data volume, it indicates that the backup data is wrong and needs to be hot-backed up again.

[0009] In another possible implementation, the data hot backup method of the present application further includes: when the computing device does not read the backup data from the cloud hard disk for a period exceeding a preset period, it indicates that the backup has ended, and the data pipeline is closed, which can save bandwidth resource overhead. Before closing the data pipeline, the heartbeat data can be stopped from being generated.

[0010] In a possible implementation, the heartbeat data block is the smallest transmission data block configured by the data pipeline. In this case, the resources used for periodically transmitting the heartbeat data block are very small, and the possibility of affecting the database business is extremely small.

[0011] A second aspect provides a data hot backup method, which is applied to a storage device of a cloud database system. After the storage device receives the heartbeat data block and the data segment of the backup data sent by the computing device, it combines at least one heartbeat data block and multiple data segments of the backup data into a data slice, and combines all the data slices into a data stream file. The method can receive the heartbeat data block to maintain the data pipeline, and perform hot backup according to the heartbeat data block and the data segment.

[0012] In a possible implementation, the backup data includes at least one of a redo log, a system table space file, a clustered index data table, or a non-clustered index data table.

[0013] The third aspect provides a computing device, which includes an interaction module and a backup module, the interaction module is used to receive backup instructions input by a user, and the backup module is used to establish a data pipeline between the computing device and the storage device in response to the backup instructions and generate a heartbeat data block; send the heartbeat data block to the storage device through the data pipeline according to the heartbeat cycle; read the backup data from the cloud hard disk; divide the backup data into multiple data segments; and send the data segments of the backup data to the storage device through the data pipeline.

[0014] In one possible implementation, the backup module is also used to generate a backup tool log including a backup status and a data volume of the backup data; when it is detected that the backup status is a backup completion status and the data volume of the backup data is consistent with a preset backup data volume, the data pipeline is closed; when it is detected that the backup status is a backup completion status and the data volume of the backup data is inconsistent with a preset backup data volume, the backup data is read from the cloud hard disk; the backup data is divided into multiple data segments; and the data segments of the backup data are sent to the storage device through the data pipeline.

[0015] In another possible implementation, the backup module is further configured to close the data pipeline when the duration of not reading backup data from the cloud hard disk exceeds a preset duration.

[0016] The explanation of terms in the third aspect, the steps executed by each module and the technical effects can be found in the corresponding description of the first aspect.

[0017] The fourth aspect provides a storage device, which includes a communication module and a processing module. The communication module is used to receive heartbeat data blocks and data segments sent by a computing device; the processing module is used to synthesize at least one heartbeat data block and multiple data segments into data fragments; and all data fragments are synthesized into a data stream file.

[0018] For the explanation of terms in the fourth aspect, the steps executed by each module and the technical effects, please refer to the corresponding description of the second aspect.

[0019] A fifth aspect provides a cloud database system, which includes the computing device of the third aspect or any possible implementation of the third aspect and the storage device of the fourth aspect.

[0020] The sixth aspect provides a computing device cluster, which includes at least one computing device, each computing device includes a processor and a memory; the processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method described in the first aspect or any possible implementation of the first aspect, or the method described in the second aspect or any possible implementation of the second aspect.

[0021] The seventh aspect provides a computer-readable storage medium, which includes computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster executes the method described in the first aspect or any possible implementation of the first aspect, or the method described in the second aspect or any possible implementation of the second aspect.

[0022] The eighth aspect provides a computer program product comprising instructions, which, when executed by a computing device cluster, causes the computing device cluster to execute the method described in the first aspect or any possible implementation of the first aspect, or the method described in the second aspect or any possible implementation of the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 A schematic diagram of hot backup of data in an existing cloud database;

[0024] Figure 2 A schematic diagram of a cloud database system in an embodiment of the present application;

[0025] Figure 3 A structural diagram of a computing device in an embodiment of the present application;

[0026] Figure 4 A structural diagram of a storage device in an embodiment of the present application;

[0027] Figure 5 A flow chart of a data hot backup method in an embodiment of the present application;

[0028] Fig. 6A A schematic diagram of sending heartbeat data blocks and data segments in an embodiment of the present application;

[0029] Figure 6B Another schematic diagram of sending heartbeat data blocks and data segments in an embodiment of the present application;

[0030] Figure 7 Another schematic diagram of the data hot backup method in the embodiment of the present application;

[0031] Figure 8 A structural diagram of a computing device in an embodiment of the present application;

[0032] Fig. 9 A structural diagram of a computing device cluster in an embodiment of the present application;

[0033] Fig.10 This is another structural diagram of a computing device cluster in an embodiment of the present application. DETAILED DESCRIPTION

[0034] Some terms involved in this application are introduced below.

[0035] Relationship database service (RDS) is also called cloud database. A database instance is the management unit of RDS. The computing device that runs the database instance includes but is not limited to a cloud server. A database instance represents an independently running RDS. Users can create and manage multiple databases in a database instance, and can access them using the same tools and applications as those for independently accessing database instances. The RDS service has no limit on the number of running instances, but each database instance has a unique identifier.

[0036] Object storage service (OBS) is an object-based storage service.

[0037] Cloud hard disk is also called elastic volume service (EVS), which can provide high-performance and high-reliability data block storage for cloud servers and support multiple performance and capacity options. Users can use EVS to dynamically increase or decrease storage capacity without downtime for maintenance. At the same time, EVS also supports data snapshot and data replication functions, which can effectively protect user data security and meet disaster recovery and disaster recovery needs.

[0038] Hot backup is a backup method when the database is in normal operation. Hot backup will not interrupt business. The database logs involved in hot backup include change logs and redo logs. Change logs include but are not limited to binary log files. Binary log files are logical logs, usually referred to as binlogs, which record statements for updating the database. Database statements include data definition language (DDL) and data manipulation language (DML) statements. Binary log files are stored in the disk in binary form, but do not contain statements that do not modify any data, such as data query statements (such as select statements, show statements), etc. Redo logs are physical logs used to record the content of the data page update. Redo logs include transaction identifiers, operation types (insert, update, delete), operation objects (table names, row numbers), data values ​​before and after the operation, etc. This information can be used to restore all committed transactions in the event of an abnormal situation, and to make the restored database consistent with the database before the downtime.

[0039] exist Figure 1In the hot backup process shown, since the backup data is sent in the form of a data stream, and the cloud server can only send it when resources are idle, the data stream of the backup data is easily interrupted during the process of transferring the backup data from the cloud hard disk 1 to the storage device. After the cloud server divides the backup data into data slices, this method uses the cloud hard disk 2 to store the data slices of the backup data. In the process of cloud hard disk 2 sending the data slices to the storage device, there is no need to wait for the cloud server to generate the data slices, thereby reducing the backup interruption. However, cloud hard disk 2 needs to occupy storage resources to store data slices, and the cloud server needs to occupy network bandwidth resources to send the data slices of the backup data to the cloud hard disk. It also requires computing resources for the cloud server to read and write data. Therefore, the above hot backup method requires a lot of resource overhead.

[0040] The hot backup method of the present application can use heartbeat data to keep the data pipeline connected, so that there is no need to read and write data from the cloud hard disk 2, thereby reducing the resource overhead of hot backup. The data hot backup method of the present application can be applied to a cloud database system, see Figure 2 In one embodiment, the cloud database system includes a computing device 100 , a cloud hard disk 200 , and a storage device 300 .

[0041] The computing device 100 is used to receive a backup instruction input by a user; establish a data pipeline with the storage device in response to the backup instruction and generate a heartbeat data block; send the heartbeat data block to the storage device through the data pipeline according to the heartbeat cycle; read the backup data from the cloud hard disk 200, and divide the backup data into multiple data segments; send the data segments of the backup data to the storage device 300 through the data pipeline;

[0042] The storage device 300 is used to receive the heartbeat data blocks and data segments sent by the computing device; combine one or more heartbeat data blocks and multiple data segments into data slices; and combine all data slices into a data stream file.

[0043] The computing device 100, the cloud hard disk 200, and the storage device 300 can all be implemented by software or by hardware. As an example, the implementation of the computing device 100 is described below. Similarly, the implementation of the cloud hard disk 200 and the storage device 300 can refer to the implementation of the computing device 100.

[0044] As an example of a software functional unit, the computing device 100 may include code running on a computing instance. Among them, the computing instance may be at least one of a physical host (computing device), a virtual machine, a container and other computing devices. Furthermore, the above-mentioned computing device may be one or more. For example, the computing device 100 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the application may be distributed in the same region or in different regions. The multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple data centers with close geographical locations. Among them, usually a region may include multiple AZs.

[0045] Similarly, multiple hosts / virtual machines / containers used to run the code can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, a VPC is set up in a region. For cross-region communication between two VPCs in the same region and between VPCs in different regions, a communication gateway must be set up in each VPC to achieve interconnection between VPCs through the communication gateway.

[0046] As an example of a hardware functional unit, the computing device 100 may include at least one computing device, such as a server, etc. Alternatively, the computing device 100 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0047] The multiple computing devices included in the computing device 100 can be distributed in the same region or in different regions. The multiple computing devices included in the computing device 100 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the computing device 100 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0048] The following is an introduction to the computing device, cloud hard disk and storage device for implementing the data hot backup method in this application. Figure 3 In one embodiment, the present application provides a computing device 100 including an interaction module 101 and a backup module 102, wherein the interaction module 101 is used to receive a backup instruction input by a user; the backup module 102 is used to establish a data pipeline between the computing device 100 and the storage device 300 and generate a heartbeat data block in response to the backup instruction; send the heartbeat data block to the storage device 300 through the data pipeline according to the heartbeat cycle; read the backup data from the cloud hard disk 200, and divide the backup data into multiple data segments; and send the data segments of the backup data to the storage device 300 through the data pipeline.

[0049] The interaction module 101 and the backup module 102 can be implemented by software or hardware. For example, the backup module 102 is taken as an example to introduce the implementation of the backup module 102. Similarly, the implementation of the interaction module 101 can refer to the implementation of the backup module 102.

[0050] As an example of a software functional unit, the backup module 102 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, the backup module 102 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same AZ or in different AZs, each AZ including one data center or multiple data centers with close geographical locations. Generally, a region may include multiple AZs.

[0051] Similarly, multiple hosts / virtual machines / containers used to run the code can be distributed in the same VPC or in multiple VPCs. Usually, a VPC is set up in a region. For cross-region communication between two VPCs in the same region and between VPCs in different regions, a communication gateway needs to be set up in each VPC to achieve interconnection between VPCs through the communication gateway.

[0052] As an example of a hardware functional unit, the backup module 102 may include at least one computing device, such as a server, etc. Alternatively, the backup module 102 may also be a device implemented using ASIC or PLD, etc. The PLD may be implemented using CPLD, FPGA, GAL or any combination thereof.

[0053] The multiple computing devices included in the backup module 102 can be distributed in the same region or in different regions. The multiple computing devices included in the backup module 102 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the backup module 102 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0054] It should be noted that, in other embodiments, the backup module 102 can be used to execute any step in the data hot backup method, and the interaction module 101 can be used to execute any step in the data hot backup method. The steps that the interaction module 101 and the backup module 102 are responsible for implementing can be specified as needed. The full functions of the computing device 100 are realized by respectively implementing different steps in the data hot backup method through the interaction module 101 and the backup module 102.

[0055] The cloud hard disk 200 is used to provide cloud hard disk services, and can specifically be a storage node in a storage cluster. The architecture of the storage cluster can include but is not limited to a storage cluster of a network attached storage (NAS) architecture. The functions of the computing device 100 and the functions of the cloud hard disk 200 can be implemented by the same computing node or by distributed computing nodes.

[0056] The present application also provides a storage device 300, such as Figure 4As shown, the storage device 300 includes a communication module 301 and a processing module 302. The communication module 301 is used to receive the heartbeat data blocks and data segments sent by the computing device; the processing module 302 is used to synthesize one or more heartbeat data blocks and multiple data segments into data slices; and all data slices are synthesized into a data stream file. The communication module 301 can be a transceiver module such as a network interface card or a transceiver. The processing module 302 can include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP) or a digital signal processor (DSP). The storage device 300 can be deployed with but not limited to object storage services.

[0057] The data hot backup method of the present application is introduced below in combination with the computing device, cloud hard disk and storage device of the present application. Figure 5 In one embodiment, the data hot backup method of the present application includes:

[0058] Step 501: The computing device 100 receives a backup instruction input by a user.

[0059] When the computing device 100 is in a state of running a database instance, the backup instruction received by the computing device 100 is used to instruct the computing device 100 to perform a hot backup. During the hot backup process, the database can still be used normally.

[0060] Step 502: The computing device 100 establishes a data pipeline and generates a heartbeat data block in response to a backup instruction.

[0061] Optionally, after receiving the backup instruction, the computing device 100 sends a request to establish a hypertext transfer protocol secure (HTTPS) connection to the storage device 300 according to the backup instruction. After the storage device 300 feeds back confirmation information based on the request to establish the HTTPS connection, the computing device 100 establishes a data pipeline between the computing device 100 and the storage device 300.

[0062] The data pipeline of the present application is used to persistently store the data continuously generated by the computing device 100 in the database of the storage device 300. Optionally, the data pipeline is a stream data pipeline between the computing device 100 and the storage device 300, which is used to transmit the data stream. After the data is input into the data pipeline, the data pipeline can perform data conversion so that the data format meets the storage requirements of the database.

[0063] In the present application, the content of the heartbeat data block may be pre-configured, and may specifically include but is not limited to the identification of the heartbeat data block. The heartbeat data block is smaller than the data slice, for example, the heartbeat data block is smaller than 1 megabit. The size of the data slice may be several hundred megabytes (MB) or several gigabytes (GB), usually a value between 128MB and 5GB.

[0064] Optionally, the heartbeat data block can be the smallest data block allowed to be transmitted by the data pipeline, such as 64KB (kilobytes). Since the heartbeat data block is sent continuously, the amount of data transmitted is small, and the impact on the database business of the computing device 100 is small. The size of the heartbeat data block of the present application can also be set to other values, such as 128KB, 256KB, 512KB, etc., without specific limitation. The amount of data in the heartbeat data block is one or several orders of magnitude smaller than the data fragmentation, so the data pipeline between the computing device 100 and the storage device 300 can be maintained by transmitting a small amount of data.

[0065] Step 503: The computing device 100 sends the heartbeat data block to the storage device 300 through the data pipeline according to the heartbeat cycle.

[0066] A heartbeat data block refers to a data block sent according to a heartbeat cycle, and the heartbeat cycle is less than the connection duration threshold of the data pipeline. Within the value range of the heartbeat cycle, the larger the heartbeat cycle, the lower the frequency of sending heartbeat data blocks, so the fewer network resources are required. For example, the heartbeat cycle is 1 second less than the connection duration threshold of the data pipeline, so sending a small amount of heartbeat data blocks can keep the data pipeline connected. The heartbeat cycle can be any value from 1 to 30 seconds, such as 5 seconds, 10 seconds, 20 seconds, or it can be an empirical value set according to actual conditions. The heartbeat cycle and the connection duration threshold of the data pipeline can be set according to actual conditions, and this application is not limited thereto.

[0067] Step 504 , the computing device 100 reads the backup data from the cloud hard disk 200 .

[0068] Specifically, the computing device 100 sends a read data instruction to the cloud hard disk 200, and the read data instruction carries the address of the backup data; the cloud hard disk 200 obtains the backup data according to the address carried by the read data instruction, and sends the backup data to the computing device 100. The backup data includes at least one item of tablespace data or redo log. In the present application, the tablespace data may include system tablespace files, clustered index data tables, and non-clustered index data tables. System tablespace files may include, but are not limited to, ibdata files. Ibdata files may include, but are not limited to, metadata of innodb tables, undo logs, the changebuffer, and the double write buffer. Clustered index data tables may include, but are not limited to innodb tables. Non-clustered index data tables refer to data tables other than innodb tables, such as auxiliary index tables.

[0069] Step 505: The computing device 100 divides the backup data into multiple data segments.

[0070] A data segment is data with a smaller granularity than a data fragment. The data segment size can be any multiple of bytes between 64KB and 10MB.

[0071] Step 506: The computing device 100 sends the data segment to the storage device 300 through the data pipeline.

[0072] Combine the following Fig. 6A and Figure 6B The process of sending heartbeat data blocks and data segments through the data pipeline is described in an example. Fig. 6A After the computing device 100 reads the backup data from the cloud hard disk 200, it divides the backup data into multiple data segments and sends the heartbeat data blocks and data segments to the storage device 300 through the data pipeline.

[0073] See also Figure 6B In one example, in heartbeat cycle 1, the computing device 100 does not obtain backup data when the table is opened after the backup is started, and the computing device 100 sends a heartbeat data block 601 to the storage device 300, but does not send a data segment 602. In heartbeat cycle 2 to heartbeat cycle n, the computing device 100 reads the backup data from the cloud hard disk 200, divides the backup data into multiple data segments 602, and sends the heartbeat data block 601 and the data segment 602 to the storage device 300 until the backup data is sent. Figure 6B The number of data segments sent in each heartbeat cycle is a schematic example and can be set according to actual conditions, and this application does not limit it.

[0074] Step 507: The storage device 300 combines one or more heartbeat data blocks and multiple data segments into data slices.

[0075] Optionally, the size of the data shard can be 1 / n of the amount of backup data, where n is a positive integer. The amount of backup data is 10 times the size of the data shard. For example, if the size of backup data is 5 GB, the size of the data shard is 500 MB. The size of the data shard can also be set to other values ​​according to actual conditions, which is not limited in this application.

[0076] Step 508: The storage device 300 combines all the data fragments into a data stream file. In this embodiment, the computing device 100 can continuously send heartbeat data blocks to the storage device 300, that is, the data flow in the data pipeline is continuous, so that the computing device 100 and the storage device 300 can maintain the data pipeline, which can reduce the backup failure caused by the connection interruption, thereby improving the reliability of the backup.

[0077] Secondly, this hot backup method does not require the use of a backup cloud hard disk, so there is no need to use the backup cloud hard disk to read and write data, which can reduce the resource overhead of hot backup.

[0078] It should be noted that the present application can perform data recovery based on the above-mentioned data stream file. Specifically, the storage device 300 can extract the heartbeat data block from the data stream file according to the heartbeat data identifier, and synthesize the extracted data block into a heartbeat file; extract the data segment of the backup data from the data stream file according to the backup data identifier; synthesize the data segments with the same backup data identifier into backup data, and then perform data recovery based on the backup data. For example, all data segments of the redo log are synthesized into a redo log, all data segments of the system table space file are synthesized into a system table space file, multiple data segments of the same clustered index data table are synthesized into a clustered index data table, and multiple data fragments of the same non-clustered index data table are synthesized into a non-clustered index data table. This can prevent the heartbeat data block from interfering with data recovery.

[0079] In an optional embodiment, the data hot backup method of the present application also includes: the computing device 100 generates a backup tool log including a backup status and the data volume of the backup data; when it is detected that the backup status is a backup completion status and the data volume of the backup data is consistent with the preset backup data volume, the data pipeline is closed; when it is detected that the backup status is a backup completion status and the data volume of the backup data is inconsistent with the preset backup data volume, steps 504 to 506 are executed.

[0080] In this embodiment, when it is detected that the backup state is the backup completion state and the amount of the backup data is consistent with the preset backup data amount, it can be determined that the backup data has been transmitted. When it is detected that the backup state is not the backup completion state, steps 504 to 506 can be executed to continue the backup. When it is detected that the backup state is the backup completion state and the amount of the backup data is inconsistent with the preset backup data amount, it indicates that an error may have occurred in the transmitted data and the backup data needs to be retransmitted.

[0081] In one example, the computing device 100 includes a backup subprocess and an upload subprocess. The backup subprocess generates a backup tool log including a backup status and a data volume of the backup data; when it is detected that the backup status is a backup completion status and the data volume of the backup data is consistent with a preset backup data volume, the heartbeat data generation stops. When the upload subprocess detects that the backup status is a backup completion status and the data volume of the backup data is consistent with a preset backup data volume, the upload subprocess closes the data pipeline.

[0082] Before the upload subprocess detects the backup status and the amount of backup data, the upload subprocess can obtain the duration of time when no data is read from the data pipeline. When the duration exceeds the time threshold, the upload subprocess is triggered to detect the backup status and the amount of backup data. Since the backup subprocess and the upload subprocess do not communicate with each other, the upload subprocess can determine the backup progress in time through the above judgment, and can close the data pipeline in time when the backup is completed, reducing resource overhead. The time threshold is greater than the heartbeat cycle, which can be but not limited to 10 seconds, and the time threshold can be but not limited to 30 seconds. The heartbeat cycle and the time threshold can be set according to actual conditions, and this application does not limit them.

[0083] In another optional embodiment, the data hot backup method of the present application further includes: when the duration of not reading backup data from the cloud hard disk 200 exceeds a preset duration, it indicates that the backup has been completed, and the computing device 100 closes the data pipeline, which can save bandwidth resources. The preset duration is greater than or equal to the time interval for actually generating data segments, and can be specifically set according to the time interval for actually generating data segments, and this application does not limit it.

[0084] For ease of understanding, the following uses ibdata files, innodb tables, and non-clustered index data tables as examples to introduce the data hot backup method of this application in conjunction with a specific hot backup scenario. Figure 7 In one example, after the computing device 100 starts the backup task, it allocates a backup subprocess and an upload subprocess to the backup task. The backup subprocess includes a backup table thread, a copy redo log thread, and a heartbeat thread. The copy redo log thread writes the redo log data into the data pipeline, and the heartbeat thread writes the heartbeat data into the data pipeline.

[0085] The backup table thread loads the table space, copies the ibdata file and the innodb table, writes the ibdata file data and the innodb table data into the data pipeline, and determines whether to set a table backup lock for the non-innodb table. If so, the non-innodb table is copied and written into the data pipeline, otherwise the backup ends. After copying the non-innodb table, determine whether to set a binlog backup lock. If so, obtain the binlog consistency site, and generate a change log based on the binlog consistency site, otherwise the backup ends. Then generate a backup tool log including the backup status, and determine whether there is a backup completion status based on the backup tool log. If so, the backup ends, otherwise the upload sub-thread obtains the pipeline data from the data pipeline and uploads the pipeline data to the object storage service database. The pipeline data includes heartbeat data blocks, redo log data, ibdata file data, innodb table data, and non-innodb table data.

[0086] The present application also provides a computing device 800. Figure 8 As shown, the computing device 800 includes: a bus 802, a processor 804, a memory 806, and a communication interface 808. The processor 804, the memory 806, and the communication interface 808 communicate through the bus 802. The computing device 800 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 800.

[0087] The bus 802 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 8 The bus 804 is represented by only one line, but does not mean that there is only one bus or one type of bus. The bus 804 may include a path for transmitting information between various components of the computing device 800 (eg, the memory 806, the processor 804, and the communication interface 808).

[0088] The processor 804 may include any one or more of a CPU, a GPU, an MP, or a DSP. The memory 806 may include a volatile memory, such as a random access memory (RAM). The processor 804 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0089] The memory 806 stores executable program codes, and the processor 804 executes the executable program codes to respectively implement the functions of the aforementioned interaction module 101 and backup module 102, thereby implementing the data hot backup method. That is, the memory 806 stores instructions for executing the data hot backup method.

[0090] Alternatively, the memory 806 stores executable codes, and the processor 804 executes the executable codes to respectively implement the functions of the computing device 100 and the storage device 300, thereby implementing the data hot backup method. That is, the memory 806 stores instructions for executing the data hot backup method.

[0091] The communication interface 808 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 800 and other devices or communication networks.

[0092] The embodiment of the present application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.

[0093] like Fig. 9 As shown, the computing device cluster includes at least one computing device 800. The memory 806 in one or more computing devices 800 in the computing device cluster may store the same instructions for executing the data hot backup method.

[0094] In some possible implementations, the memory 806 of one or more computing devices 800 in the computing device cluster may also store partial instructions for executing the data hot backup method. In other words, the combination of one or more computing devices 800 may jointly execute instructions for executing the data hot backup method.

[0095] It should be noted that the memory 806 in different computing devices 800 in the computing device cluster can store different instructions, which are respectively used to execute part of the functions of the cloud database system. That is, the instructions stored in the memory 806 in different computing devices 800 can implement the functions of one or more devices among the computing device 100, the cloud hard disk 200 and the storage device 300.

[0096] In some possible implementations, one or more computing devices in the computing device cluster may be connected via a network, which may be a wide area network or a local area network. Fig.10 A possible implementation is shown. Fig.10 As shown, two computing devices 800A and 800B are connected via a network. Specifically, they are connected to the network via a communication interface in each computing device. In this type of possible implementation, the memory 806 in the computing device 800A stores instructions for executing the functions of the computing device 100 and the cloud hard disk 200. At the same time, the memory 806 in the computing device 800B stores instructions for executing the functions of the storage device 300. Fig.10 The connection method between the computing device clusters shown may be based on the consideration that the hot backup method provided in the present application requires a large amount of data to be stored, and therefore it is considered that the functions implemented by the storage device 300 are handed over to the computing device 800B for execution.

[0097] It should be understood that Fig.10 The functions of the computing device 800A shown in FIG. 8 may also be completed by multiple computing devices 800. Similarly, the functions of the computing device 800B may also be completed by multiple computing devices 800.

[0098] The embodiment of the present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes a data hot backup method.

[0099] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by the computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk). The computer-readable storage medium includes instructions that instruct the computing device to execute the data hot backup method, or instruct the computing device to execute the data hot backup method.

[0100] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the protection scope of the technical solutions of the embodiments of the present invention.

Claims

1. A data hot backup method, It is characterized in that The data hot backup method is applied to a computing device of a cloud database system, wherein the computing device is in a state of running a database instance, and the method comprises: The computing device receives a backup instruction input by a user; The computing device establishes a data pipeline between the computing device and the storage device and generates a heartbeat data block in response to the backup instruction; The computing device sends the heartbeat data block to the storage device through the data pipeline according to the heartbeat cycle; The computing device reads the backup data from the cloud hard disk; The computing device divides the backup data into a plurality of data segments; The computing device sends the data segments of the backup data to the storage device through the data pipe.

2. The method according to claim 1, It is characterized in that The method further comprises: The computing device generates a backup tool log including a backup status and a data volume of the backup data; When it is detected that the backup state is a backup completion state and the amount of the backup data is consistent with a preset backup data amount, the computing device closes the data pipeline; When it is detected that the backup status is a backup completion status and the data volume of the backup data is inconsistent with the preset backup data volume, the computing device reads the backup data from the cloud hard disk; divides the backup data into multiple data segments; and sends the data segments of the backup data to the storage device through the data pipeline.

3. The method according to claim 1, It is characterized in that The method further comprises: When the duration of not reading backup data from the cloud hard disk exceeds a preset duration, the computing device closes the data pipeline.

4. The method according to any one of claims 1 to 3, It is characterized in that The heartbeat data block is the minimum transmission data block configured for the data pipeline.

5. The method according to any one of claims 1 to 3, It is characterized in that The backup data includes at least one of a redo log, a system table space file, a clustered index data table, or a non-clustered index data table.

6. A data hot backup method, It is characterized in that The data hot backup method is applied to a storage device of a cloud database system, and the method comprises: The storage device receives the heartbeat data block and the data segment of the backup data sent by the computing device; The storage device combines at least one heartbeat data block and a plurality of data segments of backup data into data slices; The storage device combines all data segments into a data stream file.

7. The method according to claim 6, It is characterized in that The backup data includes at least one of a redo log, a system table space file, a clustered index data table, or a non-clustered index data table.

8. A computing device, It is characterized in that include: An interactive module, used for receiving a backup instruction input by a user; A backup module, configured to establish a data pipeline between the computing device and the storage device and generate a heartbeat data block in response to the backup instruction; Sending the heartbeat data block to a storage device through the data pipeline according to the heartbeat cycle; Reading backup data from a cloud hard disk; dividing the backup data into multiple data segments; The data segments of the backup data are sent to the storage device through the data pipe.

9. The computing device according to claim 8, It is characterized in that The backup module is further used to generate a backup tool log including a backup status and the data volume of the backup data; when it is detected that the backup status is a backup completion status and the data volume of the backup data is consistent with a preset backup data volume, close the data pipeline; when it is detected that the backup status is a backup completion status and the data volume of the backup data is inconsistent with a preset backup data volume, read the backup data from the cloud hard disk; Dividing the backup data into a plurality of data segments; The data segments of the backup data are sent to the storage device through the data pipe.

10. The computing device according to claim 8, It is characterized in that The backup module is further configured to close the data pipeline when the duration for which no backup data is read from the cloud hard disk exceeds a preset duration.

11. A computing device according to any one of claims 8 to 10, It is characterized in that The heartbeat data block is the minimum transmission data block configured for the data pipeline.

12. A computing device according to any one of claims 8 to 10, It is characterized in that The backup data includes at least one of a redo log, a system table space file, a clustered index data table, or a non-clustered index data table.

13. A storage device, It is characterized in that include: A communication module, used for receiving the heartbeat data blocks and data segments of the backup data sent by the computing device; A processing module, used for combining at least one heartbeat data block and a plurality of data segments of backup data into data fragments; Combine all data segments into a data stream file.

14. The device according to claim 13, It is characterized in that The backup data includes at least one of a redo log, a system table space file, a clustered index data table, or a non-clustered index data table.

15. A cloud database system, It is characterized in that The method comprises a computing device as claimed in any one of claims 8 to 12 and a storage device as claimed in any one of claims 13 to 14.

16. A computing device cluster, It is characterized in that It includes at least one computing device, each computing device includes a processor and a memory; the processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method as described in any one of claims 1 to 7.

17. A computer-readable storage medium, It is characterized in that The method comprises computer program instructions, and when the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method according to any one of claims 1 to 7.

18. A computer program product comprising instructions, It is characterized in that When the instructions are executed by a computing device cluster, the computing device cluster is caused to perform the method according to any one of claims 1 to 7.